hashcodecs 1.3.0__tar.gz → 1.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/.gitignore +3 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/BENCHMARK.md +35 -3
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/CHANGELOG.md +31 -1
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/CITATION.cff +2 -2
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/Cargo.lock +1 -1
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/Cargo.toml +2 -1
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/PKG-INFO +21 -11
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/README.md +20 -10
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/SAFETY.md +9 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/xxhash.rs +74 -3
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/ARCHITECTURE.md +30 -9
- hashcodecs-1.4.0/docs/requirements.txt +2 -0
- hashcodecs-1.3.0/src/bindings/base64/schema_generated.rs → hashcodecs-1.4.0/generated/rust/binding_schema.rs +563 -51
- hashcodecs-1.4.0/generated/rust/murmur3_classes.rs +34 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/_hashcodecs.pyi +20 -17
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/pyproject.toml +21 -11
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/alphabet.rs +14 -8
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/backend.rs +1 -0
- hashcodecs-1.4.0/src/base64/decode/aarch64.rs +354 -0
- hashcodecs-1.4.0/src/base64/decode/avx2.rs +338 -0
- hashcodecs-1.4.0/src/base64/decode/avx512.rs +355 -0
- hashcodecs-1.4.0/src/base64/decode/sse41.rs +129 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/decode/ssse3.rs +99 -9
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/decode/tables.rs +26 -1
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/decode/x86_contracts.rs +37 -4
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/decode.rs +198 -21
- hashcodecs-1.4.0/src/base64/encode/aarch64.rs +274 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/encode/avx2.rs +260 -101
- hashcodecs-1.4.0/src/base64/encode/avx512.rs +309 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/encode/cache.rs +6 -0
- hashcodecs-1.4.0/src/base64/encode/ssse3.rs +165 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/encode.rs +290 -5
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/miri_tests.rs +2 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/output_buffer.rs +2 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/proofs.rs +2 -0
- hashcodecs-1.4.0/src/base64/runtime_dispatch.rs +918 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/tests/aarch64.rs +45 -7
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/tests.rs +630 -18
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64.rs +11 -5
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/bindings/arguments.rs +6 -73
- hashcodecs-1.3.0/src/bindings/base64/callbacks.rs → hashcodecs-1.4.0/src/bindings/base64/api.rs +58 -37
- hashcodecs-1.4.0/src/bindings/base64/batch.rs +450 -0
- hashcodecs-1.4.0/src/bindings/base64/configured.rs +800 -0
- hashcodecs-1.4.0/src/bindings/base64/configured_tests.rs +911 -0
- hashcodecs-1.4.0/src/bindings/base64/decode.rs +648 -0
- hashcodecs-1.4.0/src/bindings/base64/encode.rs +679 -0
- hashcodecs-1.4.0/src/bindings/base64/lenient.rs +549 -0
- hashcodecs-1.4.0/src/bindings/base64/policy.rs +463 -0
- {hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/helpers → hashcodecs-1.4.0/src/bindings/base64/scan}/aarch64.rs +44 -1
- hashcodecs-1.4.0/src/bindings/base64/scan/scalar.rs +60 -0
- {hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/helpers → hashcodecs-1.4.0/src/bindings/base64/scan}/x86.rs +86 -27
- hashcodecs-1.4.0/src/bindings/base64/scan.rs +182 -0
- hashcodecs-1.4.0/src/bindings/base64/staging.rs +380 -0
- hashcodecs-1.4.0/src/bindings/base64/strict.rs +274 -0
- hashcodecs-1.4.0/src/bindings/base64.rs +14 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/bindings/buffer.rs +44 -8
- hashcodecs-1.4.0/src/bindings/compatibility.rs +113 -0
- hashcodecs-1.4.0/src/bindings/murmur3/callbacks.rs +85 -0
- hashcodecs-1.4.0/src/bindings/murmur3/digest.rs +37 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/bindings/murmur3/incremental.rs +6 -37
- hashcodecs-1.4.0/src/bindings/murmur3/methods.rs +20 -0
- hashcodecs-1.3.0/src/bindings/murmur3/mod.rs → hashcodecs-1.4.0/src/bindings/murmur3.rs +1 -1
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/bindings/objects.rs +41 -54
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/bindings/runtime.rs +3 -1
- {hashcodecs-1.3.0/src/bindings/base64 → hashcodecs-1.4.0/src/bindings}/schema.rs +20 -9
- hashcodecs-1.4.0/src/bindings/xxhash/batch.rs +472 -0
- hashcodecs-1.4.0/src/bindings/xxhash/callbacks.rs +122 -0
- {hashcodecs-1.3.0/src/bindings/base64 → hashcodecs-1.4.0/src/bindings/xxhash}/methods.rs +2 -2
- hashcodecs-1.4.0/src/bindings/xxhash.rs +5 -0
- hashcodecs-1.3.0/src/bindings/mod.rs → hashcodecs-1.4.0/src/bindings.rs +2 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x64_128.rs +1 -7
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x86_128/x86.rs +22 -12
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x86_128.rs +1 -7
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x86_32.rs +1 -7
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/batch.rs +161 -51
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/aarch64.rs +3 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/scalar.rs +9 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/x86/avx2.rs +22 -5
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/x86/avx2_batch.rs +10 -4
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/x86/avx512.rs +7 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs/x86/ssse3.rs +8 -0
- hashcodecs-1.3.0/src/xxhash/long_inputs/x86/mod.rs → hashcodecs-1.4.0/src/xxhash/long_inputs/x86.rs +0 -2
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/long_inputs.rs +71 -107
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/one_shot.rs +5 -2
- hashcodecs-1.4.0/src/xxhash/prepared.rs +136 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/primitives.rs +35 -7
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/proofs.rs +8 -2
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/short_inputs.rs +94 -3
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/tests.rs +127 -16
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash.rs +2 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/tools/generate_api_metadata.py +76 -153
- hashcodecs-1.4.0/tools/install_local_wheel.py +30 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/tools/verify_sdist.py +1 -0
- hashcodecs-1.3.0/docs/benchmarks/base64-python-batch-large.svg +0 -85
- hashcodecs-1.3.0/docs/benchmarks/base64-python-batch-memoryview.svg +0 -238
- hashcodecs-1.3.0/docs/benchmarks/base64-python-batch-reusable.svg +0 -106
- hashcodecs-1.3.0/docs/benchmarks/base64-python-batch.svg +0 -313
- hashcodecs-1.3.0/docs/benchmarks/base64-python-lenient.svg +0 -127
- hashcodecs-1.3.0/docs/benchmarks/base64-python-memoryview.svg +0 -159
- hashcodecs-1.3.0/docs/benchmarks/base64-python-mutable.svg +0 -83
- hashcodecs-1.3.0/docs/benchmarks/base64-python-reusable.svg +0 -115
- hashcodecs-1.3.0/docs/benchmarks/base64-python.svg +0 -203
- hashcodecs-1.3.0/docs/benchmarks/base64-rust.svg +0 -203
- hashcodecs-1.3.0/docs/benchmarks/murmur3-python-mutable.svg +0 -169
- hashcodecs-1.3.0/docs/benchmarks/murmur3-python.svg +0 -235
- hashcodecs-1.3.0/docs/benchmarks/murmur3-rust.svg +0 -193
- hashcodecs-1.3.0/docs/benchmarks/performance-at-a-glance.svg +0 -37
- hashcodecs-1.3.0/docs/benchmarks/results.csv +0 -587
- hashcodecs-1.3.0/docs/benchmarks/xxh3-python.svg +0 -181
- hashcodecs-1.3.0/docs/benchmarks/xxh3-rust-batch-remainders.svg +0 -139
- hashcodecs-1.3.0/docs/benchmarks/xxh3-rust.svg +0 -169
- hashcodecs-1.3.0/docs/requirements.txt +0 -2
- hashcodecs-1.3.0/src/base64/decode/aarch64.rs +0 -230
- hashcodecs-1.3.0/src/base64/decode/avx2.rs +0 -180
- hashcodecs-1.3.0/src/base64/decode/avx512.rs +0 -125
- hashcodecs-1.3.0/src/base64/decode/sse41.rs +0 -59
- hashcodecs-1.3.0/src/base64/encode/aarch64.rs +0 -122
- hashcodecs-1.3.0/src/base64/encode/avx512.rs +0 -134
- hashcodecs-1.3.0/src/base64/encode/ssse3.rs +0 -86
- hashcodecs-1.3.0/src/base64/runtime_dispatch.rs +0 -369
- hashcodecs-1.3.0/src/bindings/base64/batch.rs +0 -236
- hashcodecs-1.3.0/src/bindings/base64/decode/batch.rs +0 -125
- hashcodecs-1.3.0/src/bindings/base64/decode/fallback.rs +0 -146
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced/config.rs +0 -136
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced/scanner.rs +0 -354
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced/specials.rs +0 -73
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced/staging.rs +0 -152
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced.rs +0 -183
- hashcodecs-1.3.0/src/bindings/base64/decode/native/advanced_tests.rs +0 -574
- hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/compat.rs +0 -18
- hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/helpers/mod.rs +0 -130
- hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/helpers/scalar.rs +0 -41
- hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/mod.rs +0 -131
- hashcodecs-1.3.0/src/bindings/base64/decode/native/lenient/state_machine.rs +0 -241
- hashcodecs-1.3.0/src/bindings/base64/decode/native/strict.rs +0 -444
- hashcodecs-1.3.0/src/bindings/base64/decode/native.rs +0 -23
- hashcodecs-1.3.0/src/bindings/base64/decode/output.rs +0 -424
- hashcodecs-1.3.0/src/bindings/base64/decode/plan.rs +0 -371
- hashcodecs-1.3.0/src/bindings/base64/decode.rs +0 -88
- hashcodecs-1.3.0/src/bindings/base64/encode/batch.rs +0 -140
- hashcodecs-1.3.0/src/bindings/base64/encode.rs +0 -260
- hashcodecs-1.3.0/src/bindings/base64/mod.rs +0 -293
- hashcodecs-1.3.0/src/bindings/murmur3/digest.rs +0 -25
- hashcodecs-1.3.0/src/bindings/murmur3/methods.rs +0 -108
- hashcodecs-1.3.0/src/bindings/murmur3/one_shot.rs +0 -86
- hashcodecs-1.3.0/src/bindings/xxhash/batch.rs +0 -374
- hashcodecs-1.3.0/src/bindings/xxhash/methods.rs +0 -205
- hashcodecs-1.3.0/src/bindings/xxhash/mod.rs +0 -195
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/LICENSE +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/LICENSE-MIT +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/SECURITY.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/Cargo.toml +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/base64.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/crossover.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/murmur3.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/benches/support/mod.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/build.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/api/base64.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/api/murmur3.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/api/xxh3.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/base64.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/compatibility.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/index.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/murmur3.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/performance.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/docs/xxh3.md +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/__init__.py +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/__init__.pyi +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/base64.py +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/base64.pyi +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/murmur3.py +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/murmur3.pyi +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/py.typed +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/xxhash.py +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hashcodecs/xxhash.pyi +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/hatch_build.py +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/backend.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/base64/error.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/lib.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/block_buffer.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/dispatch.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/miri_tests.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/primitives.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/proofs.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/tests.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x64_128/x86.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3/x86_32/x86.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/murmur3.rs +0 -0
- {hashcodecs-1.3.0 → hashcodecs-1.4.0}/src/xxhash/miri_tests.rs +0 -0
|
@@ -9,6 +9,10 @@ Build the Python wheel with CPython 3.12 and the full C API. Keep competitor val
|
|
|
9
9
|
Use `uv run --python 3.12 --no-project python benchmarks/render_charts.py` to render the charts. Read exact values in
|
|
10
10
|
[docs/benchmarks/results.csv](docs/benchmarks/results.csv).
|
|
11
11
|
|
|
12
|
+
The standard and URL-safe Python decode panels report CPython 3.12.10 measurements from 2026-09-09; the lenient
|
|
13
|
+
decode chart retains its 2026-09-05 measurements. Each value is the median of 15 samples lasting at least 0.2
|
|
14
|
+
seconds each, with one logical CPU pinned.
|
|
15
|
+
|
|
12
16
|
## Timing Controls
|
|
13
17
|
|
|
14
18
|
Every Python benchmark accepts `--samples` (default: 15) and `--minimum-sample-seconds` (default: 0.2). Their
|
|
@@ -34,15 +38,38 @@ For Python, run the upstream `xxhash` extension beside hashcodecs. Pass 32 equal
|
|
|
34
38
|
The Rust remainder cases pass two or three equal-size long inputs. Run Python remainder cases with
|
|
35
39
|
`python benchmarks/python_xxhash.py --batch-counts 2 3`.
|
|
36
40
|
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
41
|
+
To measure a branch against its parent, build both extensions with the same Python and Rust toolchains, then run:
|
|
42
|
+
|
|
43
|
+
```sh
|
|
44
|
+
uv run --frozen --no-sync python benchmarks/compare_xxh3_batches.py path/to/parent/_hashcodecs.pyd path/to/branch/_hashcodecs.pyd
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Use the corresponding `.so` paths on Linux or macOS. The comparison loads both builds into one interpreter, pins
|
|
48
|
+
one CPU, and alternates their timing order across 15 samples. It checks matching digests and covers bytes,
|
|
49
|
+
bytearrays, and writable memoryviews. Use `--batch-counts 2 9 32 33 --sizes 64` to inspect small batches and the
|
|
50
|
+
32-result stack boundary. Positive `change_percent` values mean higher branch throughput.
|
|
51
|
+
|
|
52
|
+
The [32-item parent comparison](docs/benchmarks/xxh3-batch-parent-comparison.csv) records CPython 3.12.10 results
|
|
53
|
+
against parent commit `f17ab86`, measured on 2026-09-05. These paired measurements also cover the 256 KiB
|
|
54
|
+
GIL-detachment threshold and 1 MiB items. The Python XXH3 chart uses the candidate's allocating-batch measurements;
|
|
55
|
+
packed-output and upstream values retain their prior measurements.
|
|
56
|
+
|
|
57
|
+
Across the 30 paired 32-item cases, branch throughput ranges from 2.59% lower to 2.70% higher than the parent.
|
|
58
|
+
The [stack-boundary comparison](docs/benchmarks/xxh3-batch-boundary-comparison.csv) covers counts 2 and 33 with
|
|
59
|
+
64-byte items: two-item bytearray batches lose 5.40–6.05%, and 33-item bytes batches lose 6.06–7.05%. These
|
|
60
|
+
measurements show residual overhead for some small-input batches; they do not establish zero regression.
|
|
61
|
+
|
|
62
|
+
The Rust mixed benchmarks use `[1024, 1024, 4096, 4096]`, `[257, 258, 259, 260]`, `[240, 240, 241, 241]`, and the
|
|
63
|
+
reverse boundary order. The 1024/4096 case measures adjacent two-item long runs. The 257–260 case measures a
|
|
64
|
+
four-item run with one shared stripe count and distinct final stripes. The 240/241 cases measure both orders across
|
|
65
|
+
the short/long dispatch boundary.
|
|
40
66
|
|
|
41
67
|
Use the focused one-shot run to cover the AVX2 four-chain boundaries:
|
|
42
68
|
|
|
43
69
|
```sh
|
|
44
70
|
cargo bench --manifest-path benches/Cargo.toml --bench xxhash -- "xxh3_(64|128)/(240|241|512|768|1024|1536|2048|4096)/hashcodecs"
|
|
45
71
|
cargo bench --manifest-path benches/Cargo.toml --bench xxhash -- "xxh3_batch/mixed/.*/hashcodecs_(64|128)"
|
|
72
|
+
cargo bench --manifest-path benches/Cargo.toml --bench xxhash -- "xxh3_prepared"
|
|
46
73
|
```
|
|
47
74
|
|
|
48
75
|
[](docs/benchmarks/xxh3-rust.svg)
|
|
@@ -64,6 +91,11 @@ noisy cases insert `!` at the same boundaries. Both cases measure returned bytes
|
|
|
64
91
|
|
|
65
92
|
[](docs/benchmarks/base64-python-lenient.svg)
|
|
66
93
|
|
|
94
|
+
## Wrapped Python Base64
|
|
95
|
+
|
|
96
|
+
Run `python benchmarks/python_base64.py --wrapped` with CPython 3.15 or newer. The benchmark inserts newlines after
|
|
97
|
+
76 output characters and measures returned bytes and a reusable `bytearray`.
|
|
98
|
+
|
|
67
99
|
## Python Memoryview Inputs
|
|
68
100
|
|
|
69
101
|
Use `--memoryview-input` for full immutable views and `--sliced-memoryview-input` for equal-length contiguous views
|
|
@@ -4,6 +4,35 @@ This file records notable user-facing changes to `hashcodecs`. Version 1.0.0 sta
|
|
|
4
4
|
|
|
5
5
|
## [Unreleased]
|
|
6
6
|
|
|
7
|
+
## [1.4.0] - 2026-09-10
|
|
8
|
+
|
|
9
|
+
### What's Changed
|
|
10
|
+
* refactor: flatten internal module layout by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/91
|
|
11
|
+
* refactor: prepare Base64 decoder policies once by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/92
|
|
12
|
+
* refactor: consolidate binding and dispatch cleanup by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/93
|
|
13
|
+
* refactor: simplify project module layout by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/94
|
|
14
|
+
* refactor: prepare codec policies and unify bindings by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/95
|
|
15
|
+
* refactor: accelerate Base64 codec hot paths by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/96
|
|
16
|
+
* fix: protect XXH3 batches from GC reentrancy by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/97
|
|
17
|
+
* refactor: unify Base64 decode execution by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/98
|
|
18
|
+
* refactor: flatten Base64 binding ownership by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/99
|
|
19
|
+
* refactor: improve XXH3 loads and internal naming by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/100
|
|
20
|
+
* refactor: reduce Base64 decoder setup and retry costs by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/101
|
|
21
|
+
* fix: preserve Base64 decode boundaries and restore ARM coverage by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/102
|
|
22
|
+
* refactor: keep XXH3 AVX2 tail chains in registers by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/103
|
|
23
|
+
* refactor: store AVX2 Base64 vectors in smaller groups by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/104
|
|
24
|
+
* fix: reduce Python overhead and align canonical decoding by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/106
|
|
25
|
+
* refactor: reduce cached AVX2 encoder setup by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/107
|
|
26
|
+
* feat: optimize XXH3 seeded and batch hashing by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/108
|
|
27
|
+
* Optimize AVX2 Base64 decoding by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/109
|
|
28
|
+
* refactor: optimize Python batch orchestration by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/110
|
|
29
|
+
* feat: optimize XXH3 batches and wrapped Base64 by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/111
|
|
30
|
+
* Fuse Base64 SIMD lookup paths and Murmur3 hex output by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/112
|
|
31
|
+
* fix: preserve custom alphabets and batch typing by @kozistr in https://github.com/kozistr/hashcodecs-rs/pull/113
|
|
32
|
+
|
|
33
|
+
|
|
34
|
+
**Full Changelog**: https://github.com/kozistr/hashcodecs-rs/compare/v1.3.0...v1.4.0
|
|
35
|
+
|
|
7
36
|
## [1.3.0] - 2026-09-04
|
|
8
37
|
|
|
9
38
|
### What's Changed
|
|
@@ -164,7 +193,8 @@ This file records notable user-facing changes to `hashcodecs`. Version 1.0.0 sta
|
|
|
164
193
|
- Initial Python and Rust APIs for Base64 and MurmurHash3.
|
|
165
194
|
- Runtime SIMD dispatch and platform-specific CPython wheels.
|
|
166
195
|
|
|
167
|
-
[Unreleased]: https://github.com/kozistr/hashcodecs-rs/compare/v1.
|
|
196
|
+
[Unreleased]: https://github.com/kozistr/hashcodecs-rs/compare/v1.4.0...HEAD
|
|
197
|
+
[1.4.0]: https://github.com/kozistr/hashcodecs-rs/compare/v1.3.0...v1.4.0
|
|
168
198
|
[1.3.0]: https://github.com/kozistr/hashcodecs-rs/compare/v1.2.1...v1.3.0
|
|
169
199
|
[1.2.1]: https://github.com/kozistr/hashcodecs-rs/compare/v1.2.0...v1.2.1
|
|
170
200
|
[1.2.0]: https://github.com/kozistr/hashcodecs-rs/compare/v1.1.0...v1.2.0
|
|
@@ -5,8 +5,8 @@ authors:
|
|
|
5
5
|
given-names: Hyeongchan
|
|
6
6
|
orcid: https://orcid.org/0000-0002-1729-0580
|
|
7
7
|
title: "hashcodecs: SIMD-accelerated Base64, MurmurHash3, and XXH3 for Python and Rust"
|
|
8
|
-
version: 1.
|
|
9
|
-
date-released: 2026-09-
|
|
8
|
+
version: 1.4.0
|
|
9
|
+
date-released: 2026-09-10
|
|
10
10
|
license: "MIT OR Apache-2.0"
|
|
11
11
|
repository-code: "https://github.com/kozistr/hashcodecs-rs"
|
|
12
12
|
url: "https://github.com/kozistr/hashcodecs-rs"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[package]
|
|
2
2
|
name = "hashcodecs"
|
|
3
|
-
version = "1.
|
|
3
|
+
version = "1.4.0"
|
|
4
4
|
edition = "2024"
|
|
5
5
|
rust-version = "1.89"
|
|
6
6
|
autobenches = false
|
|
@@ -11,6 +11,7 @@ keywords = ["base64", "simd", "murmur3", "xxhash", "python"]
|
|
|
11
11
|
categories = ["encoding", "algorithms"]
|
|
12
12
|
readme = "README.md"
|
|
13
13
|
include = [
|
|
14
|
+
"/generated/**",
|
|
14
15
|
"/src/**",
|
|
15
16
|
"/benches/**",
|
|
16
17
|
"/tests/sanitizers.rs",
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: hashcodecs
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.4.0
|
|
4
4
|
Summary: SIMD-accelerated Base64, MurmurHash3, and xxHash codecs
|
|
5
5
|
Project-URL: Documentation, https://hashcodecs-rs.readthedocs.io/
|
|
6
6
|
Project-URL: Repository, https://github.com/kozistr/hashcodecs-rs
|
|
@@ -56,7 +56,7 @@ Move byte-heavy work into Rust without changing your Python inputs. `hashcodecs`
|
|
|
56
56
|
|
|
57
57
|
- Base64 encode and decode with standard, URL-safe, padded, unpadded, wrapped, and canonical modes.
|
|
58
58
|
- MurmurHash3 x86-32, x86-128, and x64-128 with one-shot and incremental APIs.
|
|
59
|
-
- Bit-for-bit compatible XXH3-64 and XXH3-128 with allocating and allocation-free native batch APIs.
|
|
59
|
+
- Bit-for-bit compatible XXH3-64 and XXH3-128 with prepared seeds and allocating and allocation-free native batch APIs.
|
|
60
60
|
- Caller-managed `*_into` outputs for allocation-sensitive workloads.
|
|
61
61
|
- Runtime dispatch across AVX-512, AVX2, SSE4.1, SSSE3, NEON, and scalar implementations where applicable.
|
|
62
62
|
- Direct CPython buffer handling for `bytes`, `bytearray`, and `memoryview` inputs.
|
|
@@ -147,7 +147,18 @@ assert_eq!(
|
|
|
147
147
|
0x2d06_8005_38d3_94c2
|
|
148
148
|
);
|
|
149
149
|
|
|
150
|
+
let seeded_input = vec![7; 1024];
|
|
151
|
+
let prepared = hashcodecs::xxhash::PreparedXxh3::new(42);
|
|
152
|
+
assert_eq!(
|
|
153
|
+
prepared.hash_64(&seeded_input),
|
|
154
|
+
hashcodecs::xxhash::xxh3_64(&seeded_input, 42)
|
|
155
|
+
);
|
|
156
|
+
|
|
150
157
|
let inputs: &[&[u8]] = &[b"hello", b"world"];
|
|
158
|
+
assert_eq!(
|
|
159
|
+
prepared.hash_64_batch(inputs),
|
|
160
|
+
hashcodecs::xxhash::xxh3_64_batch(inputs, 42)
|
|
161
|
+
);
|
|
151
162
|
let mut hashes = [0_u64; 2];
|
|
152
163
|
let mut index = 0;
|
|
153
164
|
hashcodecs::xxhash::xxh3_64_batch_for_each(inputs, 0, |hash| {
|
|
@@ -242,9 +253,10 @@ cargo bench --manifest-path benches/Cargo.toml --bench xxhash
|
|
|
242
253
|
|
|
243
254
|
## Performance snapshot
|
|
244
255
|
|
|
245
|
-
In the
|
|
246
|
-
|
|
247
|
-
decode.
|
|
256
|
+
In the focused 2026-09-09 hashcodecs-only run on the benchmark host, `hashcodecs.xxh3_64` processes a 1 MiB input
|
|
257
|
+
at 91.27 GiB/s. In the full 2026-09-02 run, the Base64 batch API reaches 11.84 GiB/s for encode and 6.23 GiB/s for
|
|
258
|
+
decode with 256 B items in batches of 64. Each run pins one logical CPU and uses 15 samples with a 0.2-second
|
|
259
|
+
minimum per sample. Read the
|
|
248
260
|
[benchmark details](BENCHMARK.md) and [raw comparison results](docs/benchmarks/results.csv).
|
|
249
261
|
|
|
250
262
|
## Development
|
|
@@ -258,14 +270,12 @@ uv build
|
|
|
258
270
|
Run the primary local checks:
|
|
259
271
|
|
|
260
272
|
```sh
|
|
261
|
-
|
|
262
|
-
cargo clippy --all-targets --features python -- -D warnings
|
|
263
|
-
cargo test --features python
|
|
264
|
-
uv run --frozen --no-sync ruff check . --no-cache
|
|
265
|
-
uv run --frozen --no-sync ruff format --check .
|
|
266
|
-
uv run --frozen --no-sync pytest tests --cov=hashcodecs --cov-branch --cov-fail-under=100
|
|
273
|
+
just check
|
|
267
274
|
```
|
|
268
275
|
|
|
276
|
+
`just check` builds and reinstalls the current wheel before running Python tests. Run `just full-check` before
|
|
277
|
+
committing to also check release builds, Rust core coverage, and the extracted source distribution.
|
|
278
|
+
|
|
269
279
|
Optimized paths are also checked with differential fuzzing, Kani, strict-provenance Miri, AddressSanitizer, and
|
|
270
280
|
MemorySanitizer in CI.
|
|
271
281
|
|
|
@@ -25,7 +25,7 @@ Move byte-heavy work into Rust without changing your Python inputs. `hashcodecs`
|
|
|
25
25
|
|
|
26
26
|
- Base64 encode and decode with standard, URL-safe, padded, unpadded, wrapped, and canonical modes.
|
|
27
27
|
- MurmurHash3 x86-32, x86-128, and x64-128 with one-shot and incremental APIs.
|
|
28
|
-
- Bit-for-bit compatible XXH3-64 and XXH3-128 with allocating and allocation-free native batch APIs.
|
|
28
|
+
- Bit-for-bit compatible XXH3-64 and XXH3-128 with prepared seeds and allocating and allocation-free native batch APIs.
|
|
29
29
|
- Caller-managed `*_into` outputs for allocation-sensitive workloads.
|
|
30
30
|
- Runtime dispatch across AVX-512, AVX2, SSE4.1, SSSE3, NEON, and scalar implementations where applicable.
|
|
31
31
|
- Direct CPython buffer handling for `bytes`, `bytearray`, and `memoryview` inputs.
|
|
@@ -116,7 +116,18 @@ assert_eq!(
|
|
|
116
116
|
0x2d06_8005_38d3_94c2
|
|
117
117
|
);
|
|
118
118
|
|
|
119
|
+
let seeded_input = vec![7; 1024];
|
|
120
|
+
let prepared = hashcodecs::xxhash::PreparedXxh3::new(42);
|
|
121
|
+
assert_eq!(
|
|
122
|
+
prepared.hash_64(&seeded_input),
|
|
123
|
+
hashcodecs::xxhash::xxh3_64(&seeded_input, 42)
|
|
124
|
+
);
|
|
125
|
+
|
|
119
126
|
let inputs: &[&[u8]] = &[b"hello", b"world"];
|
|
127
|
+
assert_eq!(
|
|
128
|
+
prepared.hash_64_batch(inputs),
|
|
129
|
+
hashcodecs::xxhash::xxh3_64_batch(inputs, 42)
|
|
130
|
+
);
|
|
120
131
|
let mut hashes = [0_u64; 2];
|
|
121
132
|
let mut index = 0;
|
|
122
133
|
hashcodecs::xxhash::xxh3_64_batch_for_each(inputs, 0, |hash| {
|
|
@@ -211,9 +222,10 @@ cargo bench --manifest-path benches/Cargo.toml --bench xxhash
|
|
|
211
222
|
|
|
212
223
|
## Performance snapshot
|
|
213
224
|
|
|
214
|
-
In the
|
|
215
|
-
|
|
216
|
-
decode.
|
|
225
|
+
In the focused 2026-09-09 hashcodecs-only run on the benchmark host, `hashcodecs.xxh3_64` processes a 1 MiB input
|
|
226
|
+
at 91.27 GiB/s. In the full 2026-09-02 run, the Base64 batch API reaches 11.84 GiB/s for encode and 6.23 GiB/s for
|
|
227
|
+
decode with 256 B items in batches of 64. Each run pins one logical CPU and uses 15 samples with a 0.2-second
|
|
228
|
+
minimum per sample. Read the
|
|
217
229
|
[benchmark details](BENCHMARK.md) and [raw comparison results](docs/benchmarks/results.csv).
|
|
218
230
|
|
|
219
231
|
## Development
|
|
@@ -227,14 +239,12 @@ uv build
|
|
|
227
239
|
Run the primary local checks:
|
|
228
240
|
|
|
229
241
|
```sh
|
|
230
|
-
|
|
231
|
-
cargo clippy --all-targets --features python -- -D warnings
|
|
232
|
-
cargo test --features python
|
|
233
|
-
uv run --frozen --no-sync ruff check . --no-cache
|
|
234
|
-
uv run --frozen --no-sync ruff format --check .
|
|
235
|
-
uv run --frozen --no-sync pytest tests --cov=hashcodecs --cov-branch --cov-fail-under=100
|
|
242
|
+
just check
|
|
236
243
|
```
|
|
237
244
|
|
|
245
|
+
`just check` builds and reinstalls the current wheel before running Python tests. Run `just full-check` before
|
|
246
|
+
committing to also check release builds, Rust core coverage, and the extracted source distribution.
|
|
247
|
+
|
|
238
248
|
Optimized paths are also checked with differential fuzzing, Kani, strict-provenance Miri, AddressSanitizer, and
|
|
239
249
|
MemorySanitizer in CI.
|
|
240
250
|
|
|
@@ -17,6 +17,15 @@ The checks are deliberately complementary:
|
|
|
17
17
|
exact Base64 buffers, incremental MurmurHash3 states, and XXH3 batch kernels.
|
|
18
18
|
- libFuzzer runs under sanitizers. Base64 is compared with `base64` 0.23.1,
|
|
19
19
|
MurmurHash3 with `murmur3` 0.5, and XXH3 with `xxhash-rust` 0.8.18.
|
|
20
|
+
XXH3 cases use independent equal-length and heterogeneous buffers, with
|
|
21
|
+
distinct lane contents and batch counts through nine to cover SIMD groups and tails.
|
|
22
|
+
|
|
23
|
+
The Python XXH3 batch bindings finish all input reads before allocating Python
|
|
24
|
+
result containers: a GC finalizer can clear the input list or resize a bytearray
|
|
25
|
+
even while the GIL is held. Up to 32 native results fit on the stack; larger
|
|
26
|
+
batches use a fallible vector. Large immutable inputs keep their owners across
|
|
27
|
+
GIL detachment. Subprocess tests on CPython 3.10/3.11 trigger finalizers during
|
|
28
|
+
allocation and reuse freed storage, covering both sides of the stack boundary.
|
|
20
29
|
|
|
21
30
|
Run the same checks locally on Linux:
|
|
22
31
|
|
|
@@ -6,9 +6,10 @@ use criterion::{BenchmarkId, Criterion, Throughput, criterion_group, criterion_m
|
|
|
6
6
|
|
|
7
7
|
mod support;
|
|
8
8
|
|
|
9
|
-
const SIZES: [usize;
|
|
9
|
+
const SIZES: [usize; 16] = [
|
|
10
10
|
16,
|
|
11
11
|
17,
|
|
12
|
+
32,
|
|
12
13
|
64,
|
|
13
14
|
128,
|
|
14
15
|
129,
|
|
@@ -23,8 +24,9 @@ const SIZES: [usize; 15] = [
|
|
|
23
24
|
1024 * 1024,
|
|
24
25
|
8 * 1024 * 1024,
|
|
25
26
|
];
|
|
26
|
-
const MIXED_BATCHES: [(&str, [usize; 4]);
|
|
27
|
+
const MIXED_BATCHES: [(&str, [usize; 4]); 4] = [
|
|
27
28
|
("two_long_runs", [1024, 1024, 4 * 1024, 4 * 1024]),
|
|
29
|
+
("nearby_long_lengths", [257, 258, 259, 260]),
|
|
28
30
|
("short_then_long_boundary", [240, 240, 241, 241]),
|
|
29
31
|
("long_then_short_boundary", [241, 241, 240, 240]),
|
|
30
32
|
];
|
|
@@ -140,7 +142,7 @@ fn benchmark_batch(c: &mut Criterion, group_name: &str, owned: &[Vec<u8>]) {
|
|
|
140
142
|
|
|
141
143
|
fn batch(c: &mut Criterion) {
|
|
142
144
|
for items in [2, 3, 32] {
|
|
143
|
-
for size in [64, 1024, 4 * 1024, 1024 * 1024] {
|
|
145
|
+
for size in [64, 241, 1024, 4 * 1024, 1024 * 1024] {
|
|
144
146
|
let owned = (0..items)
|
|
145
147
|
.map(|index| data(size, index as u8))
|
|
146
148
|
.collect::<Vec<_>>();
|
|
@@ -158,10 +160,79 @@ fn batch(c: &mut Criterion) {
|
|
|
158
160
|
}
|
|
159
161
|
}
|
|
160
162
|
|
|
163
|
+
fn prepared_seed(c: &mut Criterion) {
|
|
164
|
+
for size in [241, 1024] {
|
|
165
|
+
let input = data(size, 17);
|
|
166
|
+
let prepared = hashcodecs::xxhash::PreparedXxh3::new(42);
|
|
167
|
+
assert_eq!(
|
|
168
|
+
prepared.hash_64(&input),
|
|
169
|
+
hashcodecs::xxhash::xxh3_64(&input, 42)
|
|
170
|
+
);
|
|
171
|
+
assert_eq!(
|
|
172
|
+
prepared.hash_128(&input),
|
|
173
|
+
hashcodecs::xxhash::xxh3_128(&input, 42)
|
|
174
|
+
);
|
|
175
|
+
|
|
176
|
+
let mut group = c.benchmark_group(format!("xxh3_prepared/{size}"));
|
|
177
|
+
group.throughput(Throughput::Bytes(size as u64));
|
|
178
|
+
group.bench_function("one_shot_64", |bench| {
|
|
179
|
+
bench.iter(|| hashcodecs::xxhash::xxh3_64(black_box(&input), 42))
|
|
180
|
+
});
|
|
181
|
+
group.bench_function("prepared_64", |bench| {
|
|
182
|
+
bench.iter(|| prepared.hash_64(black_box(&input)))
|
|
183
|
+
});
|
|
184
|
+
group.bench_function("one_shot_128", |bench| {
|
|
185
|
+
bench.iter(|| hashcodecs::xxhash::xxh3_128(black_box(&input), 42))
|
|
186
|
+
});
|
|
187
|
+
group.bench_function("prepared_128", |bench| {
|
|
188
|
+
bench.iter(|| prepared.hash_128(black_box(&input)))
|
|
189
|
+
});
|
|
190
|
+
group.finish();
|
|
191
|
+
}
|
|
192
|
+
}
|
|
193
|
+
|
|
194
|
+
fn prepared_batch(c: &mut Criterion) {
|
|
195
|
+
for items in [2, 4] {
|
|
196
|
+
for size in [241, 1024] {
|
|
197
|
+
let owned = (0..items)
|
|
198
|
+
.map(|index| data(size, index as u8))
|
|
199
|
+
.collect::<Vec<_>>();
|
|
200
|
+
let inputs = owned.iter().map(Vec::as_slice).collect::<Vec<_>>();
|
|
201
|
+
let prepared = hashcodecs::xxhash::PreparedXxh3::new(42);
|
|
202
|
+
assert_eq!(
|
|
203
|
+
prepared.hash_64_batch(&inputs),
|
|
204
|
+
hashcodecs::xxhash::xxh3_64_batch(&inputs, 42)
|
|
205
|
+
);
|
|
206
|
+
assert_eq!(
|
|
207
|
+
prepared.hash_128_batch(&inputs),
|
|
208
|
+
hashcodecs::xxhash::xxh3_128_batch(&inputs, 42)
|
|
209
|
+
);
|
|
210
|
+
|
|
211
|
+
let mut group = c.benchmark_group(format!("xxh3_prepared_batch/{items}_items/{size}"));
|
|
212
|
+
group.throughput(Throughput::Bytes((items * size) as u64));
|
|
213
|
+
group.bench_function("batch_64", |bench| {
|
|
214
|
+
bench.iter(|| hashcodecs::xxhash::xxh3_64_batch(black_box(&inputs), 42))
|
|
215
|
+
});
|
|
216
|
+
group.bench_function("prepared_64", |bench| {
|
|
217
|
+
bench.iter(|| prepared.hash_64_batch(black_box(&inputs)))
|
|
218
|
+
});
|
|
219
|
+
group.bench_function("batch_128", |bench| {
|
|
220
|
+
bench.iter(|| hashcodecs::xxhash::xxh3_128_batch(black_box(&inputs), 42))
|
|
221
|
+
});
|
|
222
|
+
group.bench_function("prepared_128", |bench| {
|
|
223
|
+
bench.iter(|| prepared.hash_128_batch(black_box(&inputs)))
|
|
224
|
+
});
|
|
225
|
+
group.finish();
|
|
226
|
+
}
|
|
227
|
+
}
|
|
228
|
+
}
|
|
229
|
+
|
|
161
230
|
fn xxhash(c: &mut Criterion) {
|
|
162
231
|
support::pin_to_one_cpu();
|
|
163
232
|
one_shot(c);
|
|
164
233
|
batch(c);
|
|
234
|
+
prepared_seed(c);
|
|
235
|
+
prepared_batch(c);
|
|
165
236
|
}
|
|
166
237
|
|
|
167
238
|
criterion_group! {
|
|
@@ -26,13 +26,17 @@ The scalar implementations set the portability and correctness baseline. Runtime
|
|
|
26
26
|
| `src/base64.rs`, `src/base64/` | Base64 public API, operations, alphabets, output buffers, runtime dispatch, and kernels. |
|
|
27
27
|
| `src/murmur3.rs`, `src/murmur3/` | MurmurHash3 public API, variants, incremental buffers, and dispatch. |
|
|
28
28
|
| `src/xxhash.rs`, `src/xxhash/` | XXH3 public API, length-specific formulas, long-input accumulation, batching, and kernels. |
|
|
29
|
-
| `src/bindings
|
|
29
|
+
| `src/bindings.rs` | CPython extension composition root for public functions and classes. |
|
|
30
30
|
| `src/bindings/arguments.rs`, `objects.rs`, `runtime.rs` | Shared CPython parsing, object access, function registration, and GIL policy. |
|
|
31
31
|
| `src/bindings/{base64,murmur3,xxhash}/` | Algorithm-specific CPython adapters. |
|
|
32
32
|
| `hashcodecs/` | Typed Python facade and public module organization. |
|
|
33
33
|
| `benches/`, `benchmarks/` | Rust and Python throughput measurements. |
|
|
34
34
|
| `tests/`, `fuzz/` | Python compatibility tests, differential fuzzing, and safety validation. |
|
|
35
35
|
|
|
36
|
+
Layout depth is counted from the relevant source root, such as `src/`, rather than from the repository root. Build,
|
|
37
|
+
environment, cache, and generated-output directories—including `target/`, `.venv/`, `.uv-cache/`, `site/`, and
|
|
38
|
+
`__pycache__/`—are excluded from layout conventions. Checked-in generated API metadata remains under `generated/`.
|
|
39
|
+
|
|
36
40
|
## Dependency Rules
|
|
37
41
|
|
|
38
42
|
The crate groups code into layered modules within one crate:
|
|
@@ -42,7 +46,10 @@ The crate groups code into layered modules within one crate:
|
|
|
42
46
|
- dispatch modules depend on the shared CPU capability snapshot and select interchangeable kernels;
|
|
43
47
|
- architecture-specific kernels have no dependency on Python bindings;
|
|
44
48
|
- algorithm-specific Python adapters depend on the Rust APIs and shared binding policies;
|
|
45
|
-
- `bindings
|
|
49
|
+
- `bindings.rs` is a composition root and contains no parsing, buffer, or execution policy.
|
|
50
|
+
|
|
51
|
+
The core algorithms and feature-gated CPython bindings intentionally remain in this crate. Splitting them would
|
|
52
|
+
either expose private pointer-oriented implementation APIs across a crate boundary or add copies to sensitive paths.
|
|
46
53
|
|
|
47
54
|
Shared state machines own their invariants. For example, `murmur3/block_buffer.rs` keeps the pending block and its
|
|
48
55
|
length together, so each incremental hasher cannot represent an inconsistent tail. At the CPython boundary,
|
|
@@ -95,34 +102,46 @@ Output allocation and failure behavior vary by API:
|
|
|
95
102
|
- Base64 reusable-output batches stop at the first error and retain prior destination writes;
|
|
96
103
|
- the XXH3 binding validates and stabilizes all packed-batch inputs before it mutates the destination.
|
|
97
104
|
|
|
105
|
+
The Python Base64 binding stabilizes overlapping or free-threaded mutable input before decoding.
|
|
106
|
+
A prepared policy selects the same attempt order for allocating and reusable outputs. Storage adapters own
|
|
107
|
+
allocation and writes; native status values determine retries before the binding constructs Python errors.
|
|
108
|
+
Direct probes write only complete validated blocks and preserve the remaining suffix for retries.
|
|
109
|
+
Custom-alphabet probes validate the whole input first; strict attempts may write within a failing block.
|
|
110
|
+
Strict and lenient scanners keep their distinct padding transitions and share byte-processing kernels.
|
|
111
|
+
|
|
98
112
|
The Python Base64 binding sends strict input to the SIMD core without constructing a discarded Python exception.
|
|
99
113
|
For lenient input, it keeps the MIME whitespace path on normalized SIMD input and decodes other ignored bytes into
|
|
100
114
|
the final Python object or reusable buffer. The native state machine follows the padding behavior of each supported
|
|
101
115
|
CPython patch series. It calls `binascii` for malformed input that needs CPython's exact exception.
|
|
116
|
+
For older CPython padding rules, reusable-output sizing uses the decoder's SIMD alphabet-prefix scanner and
|
|
117
|
+
handles padding between runs to preserve the stop at the first complete padding sequence.
|
|
102
118
|
|
|
103
119
|
Large aligned x86 encoding may use non-temporal stores after the input exceeds the detected private-cache working
|
|
104
120
|
set. Smaller work stays on ordinary cached stores.
|
|
121
|
+
Wrapped inputs above the 4 MiB crossover send SIMD blocks to a line-aware output cursor. Blocks that fit remain
|
|
122
|
+
vector stores; the cursor splits a block into quartets only when it crosses a newline boundary.
|
|
105
123
|
|
|
106
124
|
### MurmurHash3
|
|
107
125
|
|
|
108
126
|
The `x86_32`, `x86_128`, and `x64_128` modules each own one-shot calls, incremental state, scalar block mixing,
|
|
109
|
-
tail handling, and finalization. All variants use `
|
|
127
|
+
tail handling, and finalization. All variants use `block_buffer.rs` for pending blocks and `primitives.rs` for
|
|
110
128
|
little-endian loads and finalizers. One-shot calls choose scalar, SSE4.1, or AVX2 using explicit size thresholds.
|
|
111
129
|
|
|
112
130
|
### XXH3
|
|
113
131
|
|
|
114
132
|
`one_shot.rs` selects the input-length class. `short_inputs.rs` contains the formulas for 0 to 240 bytes.
|
|
115
|
-
`long_inputs.rs` owns secret initialization, scheduling, accumulation, and merging.
|
|
116
|
-
|
|
133
|
+
`long_inputs.rs` owns secret initialization, scheduling, accumulation, and merging. `prepared.rs` retains one derived
|
|
134
|
+
secret for repeated single and batch calls with the same seed. XXH3-64 and XXH3-128 share these modules.
|
|
135
|
+
`long_inputs/aarch64.rs` and the kernels under `long_inputs/x86/` contain the ISA-specific implementations.
|
|
117
136
|
The scalar long-input flow and backend selection use the same module. These kernels handle inputs longer than 240 bytes.
|
|
118
137
|
|
|
119
138
|
The AVX2 one-shot kernel splits each full 1,024-byte block across four accumulator chains. It also splits tails
|
|
120
139
|
that contain at least four stripes, including the final overlapping stripe. The kernel reduces the chains before
|
|
121
140
|
each block scramble and before the final merge.
|
|
122
141
|
|
|
123
|
-
Native batches reuse the initialized secret and
|
|
124
|
-
inputs longer than 240 bytes use an AVX2 batch accumulator when available; single items and mixed
|
|
125
|
-
regular paths.
|
|
142
|
+
Native batches reuse the initialized secret and inspect at most four inputs before hashing the group. Two to four
|
|
143
|
+
equal-stripe-count inputs longer than 240 bytes use an AVX2 batch accumulator when available; single items and mixed
|
|
144
|
+
sizes use the regular paths.
|
|
126
145
|
Python exposes two result models:
|
|
127
146
|
|
|
128
147
|
- `xxh3_*_batch` returns `list[int]` results;
|
|
@@ -165,7 +184,9 @@ writing and preserve bytes beyond the returned length.
|
|
|
165
184
|
The typed `_hashcodecs.pyi` declaration is the canonical Python API description. It drives the public modules,
|
|
166
185
|
package exports, module stubs, native text signatures and docstrings, and API-reference member lists through
|
|
167
186
|
`tools/generate_api_metadata.py`. The generated `base64.py`, `murmur3.py`, and `xxhash.py` modules organize exports
|
|
168
|
-
without adding per-call wrappers. `
|
|
187
|
+
without adding per-call wrappers. Generated Rust schemas live under `generated/rust` and are included by thin binding
|
|
188
|
+
modules, so metadata generation never rewrites handwritten Rust source. `py.typed` makes the declarations visible to
|
|
189
|
+
type checkers.
|
|
169
190
|
|
|
170
191
|
Wheel tests execute the installed package. Coverage paths map its installed location back to the root source package.
|
|
171
192
|
|