kxco-post-quantum 1.3.0 → 1.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/BENCHMARKS.md ADDED
@@ -0,0 +1,124 @@
1
+ # Performance
2
+
3
+ Per-algorithm latency and memory, reported as tail latency rather than a mean.
4
+
5
+ A mean hides the tail, and the tail is what a request budget has to absorb. For
6
+ ML-DSA that distinction is not cosmetic: signing uses rejection sampling, so it
7
+ loops until the candidate signature falls in range, and the slow tail is inherent
8
+ to the algorithm rather than measurement noise. Anyone sizing a timeout from a
9
+ mean will size it wrong.
10
+
11
+ Reproduce with:
12
+
13
+ ```
14
+ node --expose-gc bench/primitives.mjs --iterations 100 --json bench/results/primitives.json
15
+ ```
16
+
17
+ Every parameter set the package can reach is measured, not only the five it wraps
18
+ in its own helpers.
19
+
20
+ ## Signatures
21
+
22
+ Milliseconds. 100 iterations except where the `n` column says otherwise.
23
+
24
+ | Algorithm | Operation | n | median | p95 | p99 | ops/s |
25
+ |---|---|---:|---:|---:|---:|---:|
26
+ | ML-DSA-44 | keygen | 100 | 1.695 | 2.847 | 3.317 | 545 |
27
+ | ML-DSA-44 | sign | 100 | 8.880 | 24.104 | 30.193 | 91 |
28
+ | ML-DSA-44 | verify | 100 | 1.554 | 3.000 | 3.650 | 552 |
29
+ | ML-DSA-65 | keygen | 100 | 3.394 | 9.058 | 11.699 | 238 |
30
+ | ML-DSA-65 | sign | 100 | 11.308 | 36.752 | 51.084 | 64 |
31
+ | ML-DSA-65 | verify | 100 | 3.456 | 5.032 | 5.553 | 288 |
32
+ | ML-DSA-87 | keygen | 100 | 5.226 | 8.270 | 9.635 | 194 |
33
+ | ML-DSA-87 | sign | 100 | 13.148 | 33.932 | 42.381 | 62 |
34
+ | ML-DSA-87 | verify | 100 | 4.877 | 11.119 | 21.760 | 167 |
35
+ | SLH-DSA-SHA2-128f | sign | 20 | 151.319 | 194.915 | 205.881 | 7 |
36
+ | SLH-DSA-SHA2-128f | verify | 20 | 6.412 | 10.672 | 30.298 | 121 |
37
+ | SLH-DSA-SHA2-192s | sign | 3 | 6788.860 | 7229.209 | 7229.209 | 0 |
38
+ | SLH-DSA-SHA2-192s | verify | 10 | 4.997 | 62.287 | 62.287 | 89 |
39
+ | SLH-DSA-SHAKE-256f | sign | 5 | 3333.195 | 3786.271 | 3786.271 | 0 |
40
+ | SLH-DSA-SHAKE-256f | verify | 10 | 72.219 | 86.779 | 86.779 | 14 |
41
+
42
+ **Two numbers to design around.**
43
+
44
+ **ML-DSA signing has a long tail.** ML-DSA-65 signs in 11 ms at the median and
45
+ 51 ms at p99, a factor of 4.5. Rejection sampling means an unlucky signature does
46
+ several more rounds. Size request budgets from p99, not the median.
47
+
48
+ **SLH-DSA-SHA2-192s signs in about 6.8 seconds.** That is the set `slhDsa` wraps,
49
+ and it is not a per-request operation. It suits infrequent, high-value signatures
50
+ such as firmware or root attestations. Verification is cheap, around 5 ms, so an
51
+ SLH-DSA signature is expensive to make and cheap to check. The `f` variants trade
52
+ signature size for signing speed: SHA2-128f signs in 151 ms.
53
+
54
+ ## Key encapsulation
55
+
56
+ | Algorithm | Operation | n | median | p95 | p99 | ops/s |
57
+ |---|---|---:|---:|---:|---:|---:|
58
+ | ML-KEM-512 | keygen | 100 | 0.338 | 0.969 | 1.277 | 2222 |
59
+ | ML-KEM-512 | encapsulate | 100 | 0.512 | 1.043 | 1.617 | 1720 |
60
+ | ML-KEM-512 | decapsulate | 100 | 0.845 | 2.133 | 2.561 | 965 |
61
+ | ML-KEM-768 | keygen | 100 | 0.645 | 1.357 | 1.757 | 1344 |
62
+ | ML-KEM-768 | encapsulate | 100 | 0.822 | 2.050 | 5.067 | 971 |
63
+ | ML-KEM-768 | decapsulate | 100 | 1.292 | 2.750 | 4.048 | 674 |
64
+ | ML-KEM-1024 | keygen | 100 | 1.062 | 2.043 | 2.776 | 860 |
65
+ | ML-KEM-1024 | encapsulate | 100 | 1.066 | 1.902 | 10.844 | 641 |
66
+ | ML-KEM-1024 | decapsulate | 100 | 0.917 | 1.813 | 2.169 | 918 |
67
+
68
+ ML-KEM is sub-millisecond at the median across all three sets. Moving from
69
+ Category 3 to Category 5 costs well under a millisecond per operation, so the
70
+ migration cost of ML-KEM-1024 is its 1568-byte keys and ciphertexts, not its
71
+ speed. See [MIGRATION.md](MIGRATION.md).
72
+
73
+ ## Wrapper overhead
74
+
75
+ The package helpers derive keys from a master secret through HKDF, which the raw
76
+ primitives do not. That is the only overhead they add:
77
+
78
+ | Helper | median | raw keygen median | HKDF cost |
79
+ |---|---:|---:|---:|
80
+ | `mlDsa.keypairFromMaster` | 3.945 | 3.394 | ~0.55 ms |
81
+ | `mlKem.keypairFromMaster` | 1.129 | 0.645 | ~0.48 ms |
82
+
83
+ Signing and verification go straight through, so they carry no wrapper cost
84
+ beyond hex encoding.
85
+
86
+ ## Memory
87
+
88
+ Heap growth per operation, in bytes, from `process.memoryUsage().heapUsed` across
89
+ a batch:
90
+
91
+ | Algorithm | keygen | sign or encapsulate |
92
+ |---|---:|---:|
93
+ | ML-KEM-768 | ~11.7 kB | ~13.0 kB |
94
+ | ML-KEM-1024 | ~20.6 kB | ~6.1 kB |
95
+ | ML-DSA-65 | ~12.9 kB | ~21.0 kB |
96
+ | ML-DSA-87 | ~3.3 kB | ~11.4 kB |
97
+ | SLH-DSA-SHA2-128f | ~143 kB | |
98
+ | SLH-DSA-SHA2-192s | ~192 kB | |
99
+ | SLH-DSA-SHAKE-256f | ~864 kB | |
100
+
101
+ SLH-DSA allocates one to three orders of magnitude more than the lattice schemes,
102
+ which is consistent with building a hypertree per key. The lattice figures are
103
+ tens of kilobytes and are not a constraint at these rates.
104
+
105
+ **Read these as orders of magnitude, not exact allocations.** Garbage collection
106
+ can run mid-batch, so a figure smaller than a sibling's does not reliably mean
107
+ less allocation. ML-DSA-87 keygen reading lower than ML-DSA-65 is an artefact of
108
+ collection timing, not evidence that the larger parameter set allocates less.
109
+
110
+ ## What these figures are not
111
+
112
+ - **One machine, one run.** Node v26.1.0 on win32-x64. Absolute numbers move with
113
+ hardware and runtime; the ratios between operations are the portable part.
114
+ - **Not a cross-vendor comparison.** Nothing here was measured against another
115
+ implementation, so it says nothing about how this compares to a native or
116
+ hardware-backed stack. It will be slower than both.
117
+ - **Reduced samples where marked.** SLH-DSA slow variants take 3 to 20 samples
118
+ rather than 100, because 100 signatures at 6.8 seconds each is not a benchmark,
119
+ it is an afternoon. Where n is small, p95 and p99 collapse onto the maximum and
120
+ are reported that way rather than dressed up.
121
+ - **Not a side-channel measurement.** Timing here is throughput, gathered without
122
+ any attempt to detect secret-dependent variation, and it must not be read as
123
+ evidence about constant-time behaviour. See [THREAT-MODEL.md](THREAT-MODEL.md),
124
+ which states plainly that no such property is claimed.