loaderx 1.4.2__tar.gz → 1.5.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {loaderx-1.4.2/loaderx.egg-info → loaderx-1.5.2}/PKG-INFO +153 -141
- {loaderx-1.4.2 → loaderx-1.5.2}/README.md +152 -140
- {loaderx-1.4.2 → loaderx-1.5.2}/build.zig +1 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/__init__.py +1 -1
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/zrecord.py +3 -3
- {loaderx-1.4.2 → loaderx-1.5.2/loaderx.egg-info}/PKG-INFO +153 -141
- {loaderx-1.4.2 → loaderx-1.5.2}/scripts/bench.py +14 -8
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/format.zig +50 -34
- loaderx-1.5.2/src/zrecord/storage.zig +541 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord.zig +73 -36
- {loaderx-1.4.2 → loaderx-1.5.2}/tests/test_loaderx.py +16 -5
- loaderx-1.4.2/src/zrecord/storage.zig +0 -396
- {loaderx-1.4.2 → loaderx-1.5.2}/LICENSE +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/MANIFEST.in +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/build.zig.zon +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/_lib.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/dataloader.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/dataset.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/sampler.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx/utils.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/SOURCES.txt +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/dependency_links.txt +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/requires.txt +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/loaderx.egg-info/top_level.txt +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/pyproject.toml +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/scripts/build_wheels.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/setup.cfg +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/setup.py +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/codec.zig +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zrecord/exec.zig +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zsampler.zig +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/src/zstd/c.zig +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/COPYING +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/LICENSE +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/allocations.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/bits.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/bitstream.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/compiler.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/cpu.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/debug.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/debug.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/entropy_common.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/error_private.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/error_private.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/fse.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/fse_decompress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/huf.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/mem.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/pool.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/pool.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/portability_macros.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/threading.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/threading.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/xxhash.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/xxhash.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_common.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_deps.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_internal.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/common/zstd_trace.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/clevels.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/fse_compress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/hist.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/hist.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/huf_compress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_internal.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_literals.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_literals.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_sequences.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_sequences.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_superblock.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_compress_superblock.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_cwksp.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_double_fast.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_double_fast.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_fast.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_fast.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_lazy.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_lazy.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_ldm_geartab.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_opt.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstd_opt.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstdmt_compress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/compress/zstdmt_compress.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/huf_decompress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_ddict.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_ddict.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_block.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_block.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/decompress/zstd_decompress_internal.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/cover.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/cover.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/divsufsort.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/divsufsort.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/fastcover.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/dictBuilder/zdict.c +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zdict.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zstd.h +0 -0
- {loaderx-1.4.2 → loaderx-1.5.2}/vendor/zstd/lib/zstd_errors.h +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: loaderx
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.5.2
|
|
4
4
|
Summary: A compact, high-performance persistent record store with zero-copy batch gathering and transparent per-record compression, designed for single-machine AI training and serving pipelines
|
|
5
5
|
Author-email: Ben0i0d <ben0i0d@foxmail.com>
|
|
6
6
|
License: MIT License
|
|
@@ -238,9 +238,11 @@ and how much data trains it — so a caller picks a tier, never a number:
|
|
|
238
238
|
|
|
239
239
|
The sample is the byte budget the dictionary trains on (a strided subset of the
|
|
240
240
|
records), so each tier costs the same training time whatever the record size.
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
241
|
+
At realistic image scale (768 KiB records) the tiers converge — on the earlier
|
|
242
|
+
measurement box's structured data `"balanced"` and `"max"` both gather
|
|
243
|
+
~1.6 GiB/s at a 1.74x ratio — because a dictionary is a small fraction of a
|
|
244
|
+
large frame. The tiers still matter at small record sizes, where the dict is
|
|
245
|
+
most of a record and `"max"` trades gather throughput for ratio.
|
|
244
246
|
|
|
245
247
|
`"zstd_dict"` records can only be read from a store that has the dictionary
|
|
246
248
|
(`dict.zr`). The dictionary is loaded on open and shared, lock-free, across all
|
|
@@ -383,59 +385,58 @@ For practical integration examples, please refer to the **[Data2Latent](https://
|
|
|
383
385
|
the index sampler, the record store, and the full data loader — each through the
|
|
384
386
|
binding a client actually uses, so cffi, the GIL and the NumPy allocation are all
|
|
385
387
|
inside the timings. One run is enough: this is a qualitative horizontal
|
|
386
|
-
comparison,
|
|
387
|
-
|
|
388
|
-
cache.
|
|
388
|
+
comparison, sized to a realistic image workload (625 MiB of 256 KiB records)
|
|
389
|
+
that finishes `python3 scripts/bench.py all` — which prints this machine block
|
|
390
|
+
first — in under half a minute. Warm page cache.
|
|
389
391
|
|
|
390
|
-
**Machine** — one box,
|
|
391
|
-
|
|
392
|
+
**Machine** — one box, an LXC container on a server (AMD EPYC 7303, 16 cores /
|
|
393
|
+
32 threads):
|
|
392
394
|
|
|
393
395
|
| machine | value |
|
|
394
396
|
|-----------|-------|
|
|
395
|
-
| CPU | AMD
|
|
396
|
-
| frequency |
|
|
397
|
-
| caches | L1d
|
|
398
|
-
| NUMA |
|
|
399
|
-
| memory |
|
|
400
|
-
| OS | Debian GNU/Linux
|
|
401
|
-
| python | 3.14.
|
|
397
|
+
| CPU | AMD EPYC 7303 16-Core Processor, 1 socket, 16 cores / 32 threads |
|
|
398
|
+
| frequency | 1500–3437 MHz |
|
|
399
|
+
| caches | L1d 512 KiB, L1i 512 KiB, L2 8 MiB, L3 64 MiB |
|
|
400
|
+
| NUMA | 4 nodes |
|
|
401
|
+
| memory | 32 GiB (32 GiB cgroup limit) |
|
|
402
|
+
| OS | Debian GNU/Linux 13 (trixie), kernel 6.12.95+deb13-amd64 |
|
|
403
|
+
| python | 3.14.7, numpy 2.5.1 |
|
|
402
404
|
|
|
403
|
-
The container sees all
|
|
404
|
-
the ordinary page cache.
|
|
405
|
+
The container sees all 32 threads under a 32 GiB cgroup limit. The benchmark
|
|
406
|
+
runs on the ordinary page cache.
|
|
405
407
|
|
|
406
408
|
**1. Sampler** — index generation on its own, IID (with replacement), 1M index
|
|
407
409
|
space, against NumPy's modern API. The µs-scale figures fluctuate with box load;
|
|
408
|
-
the stable signal is the ~
|
|
410
|
+
the stable signal is the ~1.9x margin, roughly flat across batch sizes.
|
|
409
411
|
|
|
410
412
|
| sampler | batch | per batch | vs default_rng |
|
|
411
413
|
|-------------------|-------|-----------|----------------|
|
|
412
|
-
| numpy default_rng | 256 | 5.
|
|
413
|
-
| **zsampler** | 256 |
|
|
414
|
-
| numpy default_rng | 1024 |
|
|
415
|
-
| **zsampler** | 1024 |
|
|
416
|
-
| numpy default_rng | 8192 |
|
|
417
|
-
| **zsampler** | 8192 |
|
|
418
|
-
|
|
419
|
-
**2. Store** — Zrecord against the alternatives: random batch gather,
|
|
420
|
-
records, batch 256, structured (mildly compressible) data. A single
|
|
421
|
-
sample
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
hdf5's ~430).
|
|
427
|
-
|
|
428
|
-
Small records — 12 KiB per record, 3 MiB per batch:
|
|
414
|
+
| numpy default_rng | 256 | 5.6 µs | 1.00x |
|
|
415
|
+
| **zsampler** | 256 | 3.0 µs | **1.89x** |
|
|
416
|
+
| numpy default_rng | 1024 | 8.0 µs | 1.00x |
|
|
417
|
+
| **zsampler** | 1024 | 4.3 µs | **1.85x** |
|
|
418
|
+
| numpy default_rng | 8192 | 31.3 µs | 1.00x |
|
|
419
|
+
| **zsampler** | 8192 | 17.0 µs | **1.84x** |
|
|
420
|
+
|
|
421
|
+
**2. Store** — Zrecord against the alternatives: random batch gather, 2500
|
|
422
|
+
records, batch 256, structured (mildly compressible) data. A single 256 KiB
|
|
423
|
+
sample (512×512 single-channel, a realistic image record) is the horizontal
|
|
424
|
+
comparison; other record sizes are a `--shape` away. hdf5-gzip, arrayrecord,
|
|
425
|
+
torch and grain are installed and run in the default tables.
|
|
426
|
+
|
|
427
|
+
Realistic records — 256 KiB per record, 64 MiB per batch:
|
|
429
428
|
|
|
430
429
|
| store | write | gather | on disk | ratio |
|
|
431
430
|
|-------------------|------------|-------------|---------|-------|
|
|
432
|
-
| zrecord-zstd |
|
|
433
|
-
| zrecord-zstdict |
|
|
434
|
-
| zrecord-raw |
|
|
435
|
-
| npy-mmap |
|
|
436
|
-
| hdf5 |
|
|
437
|
-
|
|
|
438
|
-
|
|
|
431
|
+
| zrecord-zstd | 772 MiB/s | 4031 MiB/s | 555 MiB | 1.13x |
|
|
432
|
+
| zrecord-zstdict | 72 MiB/s | 2123 MiB/s | 388 MiB | 1.61x |
|
|
433
|
+
| zrecord-raw | 880 MiB/s | 5214 MiB/s | 625 MiB | 1.00x |
|
|
434
|
+
| npy-mmap | 1398 MiB/s | 1853 MiB/s | 625 MiB | 1.00x |
|
|
435
|
+
| hdf5 | 975 MiB/s | 1046 MiB/s | 625 MiB | 1.00x |
|
|
436
|
+
| hdf5-gzip | 37 MiB/s | 182 MiB/s | 445 MiB | 1.40x |
|
|
437
|
+
| blosc2 | 48 MiB/s | 353 MiB/s | 551 MiB | 1.13x |
|
|
438
|
+
| tensorstore | 720 MiB/s | 109 MiB/s | 462 MiB | 1.35x |
|
|
439
|
+
| arrayrecord | 331 MiB/s | 865 MiB/s | 567 MiB | 1.10x |
|
|
439
440
|
|
|
440
441
|
`write` is page-cache ingestion — the store is built and left open, exactly
|
|
441
442
|
like `np.save` and the other backends, which defer durability to the kernel.
|
|
@@ -448,34 +449,36 @@ lazy ones.
|
|
|
448
449
|
|
|
449
450
|
`gather` is fully page-cache-warmed — the benchmark sweeps every record once
|
|
450
451
|
before timing, so it measures the pure access path (memory), not first-touch
|
|
451
|
-
page faults or disk.
|
|
452
|
-
|
|
453
|
-
|
|
454
|
-
leads by ~
|
|
455
|
-
|
|
456
|
-
|
|
452
|
+
page faults or disk. At 256 KiB records the batch (64 MiB) exceeds L3, so these
|
|
453
|
+
numbers are DRAM-bandwidth-bound for every backend; the 12 KiB table that
|
|
454
|
+
predates the record-size bump was an L3-fit artifact. Fully warm, zrecord-raw
|
|
455
|
+
leads npy-mmap by ~2.8x on this box — its gather fans out across every core
|
|
456
|
+
while numpy's fancy index is single-threaded — and the gap widens where
|
|
457
|
+
single-threaded copy is slower.
|
|
457
458
|
|
|
458
459
|
**3. Loader** — the full input pipeline end to end (sample, fetch, collate, hand
|
|
459
|
-
over a batch), same workload, 4 workers. The `peak RSS`
|
|
460
|
-
uniform-codec rewrite) is the highest resident set size
|
|
461
|
-
tree while batches are flowing
|
|
462
|
-
|
|
463
|
-
|
|
464
|
-
for this workload — and it lands ~50 batches/s, an order of magnitude below the
|
|
465
|
-
rest. `--only grain` reaches it when installed.)
|
|
460
|
+
over a batch), same workload, 4 workers, batch 256 (64 MiB). The `peak RSS`
|
|
461
|
+
column (added with the uniform-codec rewrite) is the highest resident set size
|
|
462
|
+
of the whole process tree while batches are flowing. torch, grain and
|
|
463
|
+
arrayrecord are installed and wired into the defaults, and all four loaders run
|
|
464
|
+
at the full batch.
|
|
466
465
|
|
|
467
466
|
| loader | batches/s | peak RSS |
|
|
468
467
|
|----------------|-----------|----------|
|
|
469
|
-
| **loaderx** |
|
|
470
|
-
| loaderx-raw |
|
|
471
|
-
| torch |
|
|
472
|
-
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
memory
|
|
478
|
-
|
|
468
|
+
| **loaderx** | 59.2 | 2477 MiB |
|
|
469
|
+
| loaderx-raw | 80.1 | 2478 MiB |
|
|
470
|
+
| torch | 37.6 | 14525 MiB |
|
|
471
|
+
| grain | 15.8 | 4567 MiB |
|
|
472
|
+
|
|
473
|
+
At 64 MiB per batch the loader is DRAM-bound, not sampler-bound — the per-batch
|
|
474
|
+
gather (zstd ~4.0 GiB/s, raw ~5.2 GiB/s) is the whole story, and the prefetch
|
|
475
|
+
threads keep it at store-gather speed while the Python side collates. The
|
|
476
|
+
memory is the prefetch buffers plus the two dataset handles (raw + label stores
|
|
477
|
+
over the same backing data). loaderx prefetches in threads inside one process,
|
|
478
|
+
so workers share one interpreter, one numpy and one set of gather buffers; torch
|
|
479
|
+
and grain run a worker process per prefetch thread, which is most of their RSS
|
|
480
|
+
(torch's peak also includes the shared-memory collated batches). grain's
|
|
481
|
+
batches/s is the pipeline's floor here, ~5x below loaderx-raw.
|
|
479
482
|
|
|
480
483
|
**Free-threaded Python.** loaderx targets free-threaded builds (no GIL), and
|
|
481
484
|
the key sections are also measured on 3.14t — the question is whether anything
|
|
@@ -485,20 +488,22 @@ and `--transform augment`); they are not part of the default `all`:
|
|
|
485
488
|
* **Sampler** — zsampler draws ~1.9–2x faster than free-threaded numpy, the
|
|
486
489
|
same margin as on the GIL build: a batch is one Zig call either way.
|
|
487
490
|
* **Store** — the free-threaded numbers track the GIL table within ~10%.
|
|
488
|
-
`zrecord-raw` gathers
|
|
489
|
-
`zrecord-zstd`
|
|
490
|
-
build: blosc2's free-threaded wheel re-enables the GIL to load its
|
|
491
|
-
so it runs exactly as on the GIL build. Random gather,
|
|
491
|
+
`zrecord-raw` gathers 4.9 vs 5.2 GiB/s (still ~2.6x `npy-mmap`),
|
|
492
|
+
`zrecord-zstd` 3.9 vs 4.0 GiB/s. Notably, the alternatives do not gain from
|
|
493
|
+
the build: blosc2's free-threaded wheel re-enables the GIL to load its
|
|
494
|
+
extension, so it runs exactly as on the GIL build. Random gather, 256 KiB
|
|
495
|
+
records:
|
|
492
496
|
|
|
493
497
|
| store | CPython 3.14 (GIL) | free-threaded 3.14t |
|
|
494
498
|
|--------------|--------------------|---------------------|
|
|
495
|
-
| zrecord-raw |
|
|
496
|
-
| zrecord-zstd |
|
|
497
|
-
| npy-mmap |
|
|
498
|
-
| hdf5 |
|
|
499
|
-
| blosc2 |
|
|
500
|
-
| tensorstore |
|
|
501
|
-
* **Loader, identity** — unchanged, ~
|
|
499
|
+
| zrecord-raw | 5214 MiB/s | 4886 MiB/s |
|
|
500
|
+
| zrecord-zstd | 4031 MiB/s | 3913 MiB/s |
|
|
501
|
+
| npy-mmap | 1853 MiB/s | 1861 MiB/s |
|
|
502
|
+
| hdf5 | 1046 MiB/s | 1030 MiB/s |
|
|
503
|
+
| blosc2 | 353 MiB/s | 334 MiB/s |
|
|
504
|
+
| tensorstore | 109 MiB/s | 110 MiB/s |
|
|
505
|
+
* **Loader, identity** — unchanged, ~60 batches/s (loaderx) on both
|
|
506
|
+
interpreters at 64 MiB per batch.
|
|
502
507
|
* **Loader, CPU-heavy transform** — the one place the free-threaded build
|
|
503
508
|
matters. A transform runs on the prefetch threads, so the GIL serializes it on
|
|
504
509
|
standard CPython — the case where worker processes win, and the reason for the
|
|
@@ -508,13 +513,14 @@ and `--transform augment`); they are not part of the default `all`:
|
|
|
508
513
|
|
|
509
514
|
| loader | CPython 3.14 (GIL) | free-threaded 3.14t |
|
|
510
515
|
|-----------------|--------------------|---------------------|
|
|
511
|
-
| **loaderx** |
|
|
512
|
-
| loaderx-raw |
|
|
513
|
-
| torch | 408.7 | n/a |
|
|
516
|
+
| **loaderx** | 24.3 | 41.1 batches/s |
|
|
517
|
+
| loaderx-raw | 22.3 | 43.3 batches/s |
|
|
514
518
|
|
|
515
|
-
|
|
516
|
-
~1
|
|
517
|
-
|
|
519
|
+
Free-threaded still wins — the transform parallelizes across the prefetch
|
|
520
|
+
threads — ~1.7x on loaderx and ~1.9x on raw at 256 KiB. The margin is
|
|
521
|
+
narrower than at 12 KiB records (where a batch fit in cache and the transform
|
|
522
|
+
was the whole cost) because the 64 MiB gather is DRAM-bound, but wider than
|
|
523
|
+
at the 768 KiB scale, where the pipeline was even more DRAM-bound.
|
|
518
524
|
|
|
519
525
|
**Conclusion** — why the numbers look like this.
|
|
520
526
|
|
|
@@ -522,15 +528,15 @@ and `--transform augment`); they are not part of the default `all`:
|
|
|
522
528
|
zrecord gathers a whole batch and decompresses it, in a single cffi call into
|
|
523
529
|
Zig with the GIL released and the batch copied straight into its destination
|
|
524
530
|
buffer. The contenders do the same work one record at a time from Python. That
|
|
525
|
-
one fact runs through all three tables: the sampler's margin is
|
|
526
|
-
|
|
531
|
+
one fact runs through all three tables: the sampler's margin is roughly flat
|
|
532
|
+
across batch sizes (a batch costs one call either way, so the per-index work is
|
|
527
533
|
what divides them), and the store gather column is where one-record-per-call
|
|
528
534
|
costs the most. The speedup is not bought with distribution shortcuts: the IID
|
|
529
535
|
draw is unbiased like NumPy's (Lemire with rejection, so uniformity costs
|
|
530
536
|
nothing over a real index space).
|
|
531
537
|
|
|
532
538
|
**The layout matches what a training loader does.** zrecord is built for random
|
|
533
|
-
record access — the
|
|
539
|
+
record access — the record table is indexed in O(1), a dense
|
|
534
540
|
store gathers at a fixed stride. Array stores are built for contiguous scans, so
|
|
535
541
|
a scattered batch — exactly what a loader reads — fights their layout. A raw
|
|
536
542
|
gather is a fan-out across every core, where NumPy's fancy index is one thread;
|
|
@@ -540,17 +546,17 @@ on the multi-core memory subsystem that closes the gap to an uncompressed
|
|
|
540
546
|
**Compression is in the kernel, and there is one codec.** `zrecord-zstd` is not
|
|
541
547
|
"storage plus a codec": the layout, the multi-core decompress and the GIL-free
|
|
542
548
|
copy are one path, so turning compression on costs part of a margin, not an
|
|
543
|
-
order of magnitude — ~
|
|
544
|
-
|
|
545
|
-
|
|
546
|
-
|
|
547
|
-
tier) 16.5x.
|
|
549
|
+
order of magnitude — ~3.9 GiB/s here, ~11x ahead of the other compressed stores
|
|
550
|
+
(109–353 MiB/s) at the same ~1.1x ratio. The modest ratio is the data, not the
|
|
551
|
+
format: these samples are barely compressible, and on a smooth image set plain
|
|
552
|
+
zstd reaches 7.6x and the dictionary (at the `"max"` tier) 16.5x.
|
|
548
553
|
|
|
549
554
|
**The loader gap is architecture, not storage.** loaderx uses threads and never
|
|
550
555
|
ends an epoch, so a step pays no IPC and never waits on an epoch boundary; torch
|
|
551
|
-
restarts per epoch with worker processes. That is most of the ~
|
|
552
|
-
Storage also differs per loader — each reads from
|
|
553
|
-
|
|
556
|
+
restarts per epoch with worker processes. That is most of the ~1.6x over torch
|
|
557
|
+
and ~3.8x over grain here. Storage also differs per loader — each reads from
|
|
558
|
+
what it was built for — so the loader table is a different comparison than the
|
|
559
|
+
store table, not a rerun of it.
|
|
554
560
|
The one crack in the thread model is a CPU-heavy transform, which the GIL
|
|
555
561
|
serializes — that is exactly what free-threaded Python removes, so loaderx is
|
|
556
562
|
developed and benchmarked against free-threaded builds first.
|
|
@@ -563,7 +569,7 @@ page-cache ingestion with durability deferred, matching the other backends —
|
|
|
563
569
|
zrecord's own durability (`sync`/`close`) is a separate, explicit cost that
|
|
564
570
|
neither this table nor the competition pays. This is a single qualitative pass —
|
|
565
571
|
the µs-scale sampler timings and the loader `batches/s` fluctuate with box load
|
|
566
|
-
(on these
|
|
572
|
+
(on these 16 cores the compressed loader trails the raw one by ~25–30%), so
|
|
567
573
|
treat the absolute numbers as ballpark and the cross-backend margins as the
|
|
568
574
|
signal.
|
|
569
575
|
|
|
@@ -621,11 +627,15 @@ by record against the npy ground truth.
|
|
|
621
627
|
|
|
622
628
|
## Current Limitations
|
|
623
629
|
* Single-host only; multi-host training is not supported.
|
|
624
|
-
* A single sample must be at most
|
|
625
|
-
|
|
626
|
-
|
|
627
|
-
|
|
628
|
-
|
|
630
|
+
* A single sample must be at most 2 GiB (2^31 bytes). There is no fixed record
|
|
631
|
+
count: `length` is a u64 and the record table grows on demand, so how many
|
|
632
|
+
records a store holds is bounded by its total chunk capacity (up to 2^64
|
|
633
|
+
bytes) divided by the average record size — e.g. roughly 2^44 records at
|
|
634
|
+
1 MiB each, 2^33 at 2 GiB each.
|
|
635
|
+
* Metadata is read and written as the host's struct layout, so a store carries
|
|
636
|
+
the host's byte order and is not portable to a machine of the opposite
|
|
637
|
+
endianness. Every published platform is little-endian, so this only matters
|
|
638
|
+
if you build for
|
|
629
639
|
one yourself.
|
|
630
640
|
|
|
631
641
|
## Build
|
|
@@ -739,49 +749,57 @@ A record-based data runtime, focused on delivering extreme throughput and low la
|
|
|
739
749
|
`zstd` store never touches a dictionary even if one is present.
|
|
740
750
|
|
|
741
751
|
## Persistence format
|
|
742
|
-
Zrecord storage is metadata plus chunked data:
|
|
752
|
+
Zrecord storage is metadata plus chunked data. The extension is the type:
|
|
753
|
+
`.zr` files are store-global singletons, `.loc` files are record-table segments,
|
|
754
|
+
`.chunk` files are record data:
|
|
743
755
|
```
|
|
744
756
|
zrecord/
|
|
745
|
-
├── meta.zr
|
|
746
|
-
├──
|
|
747
|
-
|
|
757
|
+
├── meta.zr header (global state)
|
|
758
|
+
├── dict.zr zstd dictionary (only in dict stores)
|
|
759
|
+
├── 0.loc record table segment [0, 2^28)
|
|
760
|
+
├── 1.loc record table segment [2^28, 2^29)
|
|
761
|
+
├── 0.chunk record data
|
|
762
|
+
└── 1.chunk
|
|
748
763
|
```
|
|
749
764
|
|
|
750
|
-
### Metadata (meta.zr)
|
|
765
|
+
### Metadata (meta.zr + {id}.loc)
|
|
751
766
|
|
|
752
|
-
|
|
753
|
-
|
|
754
|
-
|
|
755
|
-
|
|
756
|
-
code.
|
|
767
|
+
Files are read and written **positionally** — pread/pwrite at computed offsets,
|
|
768
|
+
no mmap. The header is a bit-packed struct and a `.loc` segment an array of
|
|
769
|
+
16-byte `RecordLoc`s, exactly as wide as they declare, so a location is one
|
|
770
|
+
pread/pwrite of 16 bytes at a computed offset and there is no serializer
|
|
771
|
+
anywhere in the code. Every header field is byte-aligned (no bit fields cross a
|
|
772
|
+
byte), so the packed header reads as plain memory.
|
|
757
773
|
|
|
758
|
-
The
|
|
759
|
-
|
|
760
|
-
|
|
761
|
-
|
|
774
|
+
The record table is partitioned into `{id}.loc` segments so it can **grow by
|
|
775
|
+
appending a segment** instead of reserving the maximum. The id→segment mapping
|
|
776
|
+
is pure arithmetic — `seg = idx >> 28`, `off = (idx & (2^28−1)) × 16` — so a
|
|
777
|
+
segment needs no per-record bookkeeping. A segment is immutable once published,
|
|
778
|
+
so each mapping's base never moves and lock-free readers are safe to index it.
|
|
762
779
|
|
|
763
|
-
**1. Header** —
|
|
764
|
-
the table starts at a clean 16-byte offset.
|
|
780
|
+
**1. Header** — 32 bytes, the whole of meta.zr.
|
|
765
781
|
* `magic` (`ZREC`) and `version` are ordinary fields, so opening a directory
|
|
766
782
|
that is not a Zrecord store fails immediately instead of decoding garbage.
|
|
767
783
|
* `codec` is the store's one compression method, stamped at creation and
|
|
768
784
|
immutable — there is no per-record tag anywhere.
|
|
769
|
-
* `length` is the total number of records | `tail_chunk`/`tail_offset`
|
|
770
|
-
last write position.
|
|
785
|
+
* `length` (u64) is the total number of records | `tail_chunk`/`tail_offset`
|
|
786
|
+
mark the last write position.
|
|
771
787
|
* There is no chunk count. Chunks are created in order and the frontier is
|
|
772
788
|
always in the last one, so the store holds exactly chunks `0..=tail_chunk` —
|
|
773
789
|
a count would be a second copy of that fact to keep in sync.
|
|
774
790
|
|
|
775
791
|
```zig
|
|
776
792
|
const Codec = enum(u8) { raw = 0, zstd = 1, zstdict = 2, _ };
|
|
777
|
-
const Header =
|
|
778
|
-
magic:
|
|
779
|
-
tail_chunk:
|
|
793
|
+
const Header = packed struct {
|
|
794
|
+
magic: u32, version: u8, codec: Codec,
|
|
795
|
+
tail_chunk: u32, tail_offset: u32, length: u64,
|
|
796
|
+
_reserved: u80,
|
|
780
797
|
};
|
|
781
798
|
```
|
|
782
799
|
|
|
783
|
-
**2. Record table** —
|
|
784
|
-
a physical address is what makes
|
|
800
|
+
**2. Record table** — a `.loc` segment is 2^28 entries of 16 bytes (4 GiB,
|
|
801
|
+
sparse), indexed directly. Mapping an index to a physical address is what makes
|
|
802
|
+
random access efficient.
|
|
785
803
|
* `chunk_id` is the containing chunk | `offset` is the position within it |
|
|
786
804
|
`phys_length`/`logic_length` are the stored and original sizes. The codec is
|
|
787
805
|
not here: it is the header's, so a record is stored exactly the way the store
|
|
@@ -790,26 +808,17 @@ a physical address is what makes random access efficient.
|
|
|
790
808
|
```zig
|
|
791
809
|
const RecordLoc = extern struct {
|
|
792
810
|
offset: u32, phys_length: u32, logic_length: u32,
|
|
793
|
-
chunk_id:
|
|
811
|
+
chunk_id: u32,
|
|
794
812
|
};
|
|
795
813
|
```
|
|
796
814
|
|
|
797
815
|
There is no liveness flag. Every entry below `length` is live, because deletion
|
|
798
816
|
swaps the tail into the hole rather than tombstoning.
|
|
799
817
|
|
|
800
|
-
**3.
|
|
801
|
-
|
|
802
|
-
|
|
803
|
-
|
|
804
|
-
* Derivation: N = chunks × 2^32 (chunk size) / 2^20 (max record size) ⇒ a
|
|
805
|
-
maximum length of 2^24 (16777216). Equality here is the point: the store can
|
|
806
|
-
never run out of chunks before it runs out of record slots, which is what
|
|
807
|
-
lets the table be sized statically. It is a compile-time assertion.
|
|
808
|
-
* meta.zr is a fixed 16 + 2^24 × 16 bytes (256 MiB). It is sparse — a store
|
|
809
|
-
with one record allocates 4 KiB of it — and the mapping does not prefault.
|
|
810
|
-
|
|
811
|
-
### Chunk data (x.zr)
|
|
812
|
-
Densely packed record data.
|
|
818
|
+
**3. No maximum length.** The table grows a `.loc` segment at a time and the
|
|
819
|
+
data grows a chunk at a time, so there is no static record-count cap to size
|
|
820
|
+
against. The real bounds are the field widths — 2^32 chunks of 2^32 bytes
|
|
821
|
+
(2^64 bytes total), 2^31 (2 GiB) per record — and disk.
|
|
813
822
|
|
|
814
823
|
## Executor
|
|
815
824
|
|
|
@@ -829,9 +838,11 @@ Densely packed record data.
|
|
|
829
838
|
side (executed on async threads).
|
|
830
839
|
* Committed records are immutable, so the read path is lock free; `length` is
|
|
831
840
|
published to readers through an atomic.
|
|
832
|
-
* Every record is read at the offset its table entry records — the
|
|
833
|
-
|
|
834
|
-
no batching assumptions about layout.
|
|
841
|
+
* Every record is read at the offset its table entry records — the record table
|
|
842
|
+
is addressed by pure arithmetic, so random access is one pread for the
|
|
843
|
+
location and one for the bytes, with no batching assumptions about layout.
|
|
844
|
+
Locations are read once per batch (contiguous runs in one pread). Compressed
|
|
845
|
+
records are read into a
|
|
835
846
|
per-shard staging buffer and decompressed in place into the destination.
|
|
836
847
|
|
|
837
848
|
**3. Concurrency model.** `Io.Group.async` shards work by CPU count, and shards
|
|
@@ -864,7 +875,8 @@ is reclaimed by an offline `compact` that rewrites the live records in place.
|
|
|
864
875
|
that compaction cannot remove.
|
|
865
876
|
|
|
866
877
|
**5. File access.**
|
|
867
|
-
* Metadata:
|
|
878
|
+
* Metadata: `meta.zr` (32 bytes) plus one file per `.loc` segment, each 4 GiB
|
|
879
|
+
(sparse), opened when the segment is created.
|
|
868
880
|
* Chunk data: 4 GiB files created at init, accessed concurrently through
|
|
869
881
|
`readPositionalAll`/`writePositionalAll`.
|
|
870
882
|
|