@graphty/webgpu-graph-algorithms 0.2.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +37 -3
- package/package.json +3 -3
package/README.md
CHANGED
|
@@ -280,9 +280,12 @@ power of two; `benchmarks/layout-exact.bench.ts` `exactMaxNodesFromLadder`).
|
|
|
280
280
|
|
|
281
281
|
## Performance
|
|
282
282
|
|
|
283
|
-
Regenerated from `benchmarks/results
|
|
284
|
-
recorded in `docs/decisions/G3.md` appendix A; the targets
|
|
285
|
-
recorded owner decision in
|
|
283
|
+
Regenerated from the last session of each baseline under `benchmarks/results/` (`nvidia-lovelace-driver580.json`, the
|
|
284
|
+
dev box; `gpu-linux-t4.json`, the CI lane) by the procedure recorded in `docs/decisions/G3.md` appendix A; the targets
|
|
285
|
+
are the T-table of plan section 10.4. A missed target is re-fixed by a recorded owner decision in
|
|
286
|
+
`docs/decisions/G<n>.md`, never relaxed silently.
|
|
287
|
+
|
|
288
|
+
### The dev box (nvidia-lovelace-driver580)
|
|
286
289
|
|
|
287
290
|
Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-16T02:07:45.933Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z.
|
|
288
291
|
|
|
@@ -315,6 +318,37 @@ The end-to-end run of `benchmarks/layout-run.ts --nodes 100000 --edges 1000000`
|
|
|
315
318
|
iterations, batches of 8) takes 18.971 ms per iteration on the same card, uploads and readbacks included (16.975 ms of
|
|
316
319
|
GPU time per iteration in the last batch).
|
|
317
320
|
|
|
321
|
+
### The CI lane (gpu-linux-t4)
|
|
322
|
+
|
|
323
|
+
The first run of the GPU lane (`gpu.yml`, graphty-monorepo run 35316416067, 2026-09-18) on a machine.dev T4 -- one Tesla
|
|
324
|
+
T4 (16 GB), 4 vCPU of a Xeon Platinum 8259CL, driver 580.126.20 -- wrote this baseline; `scripts/bench-compare.js` fails
|
|
325
|
+
a later run of the lane whose median exceeds 3x these figures. The T-table targets were set on the dev box; the T4 meets
|
|
326
|
+
T-4 and T-5 and misses T-1 (both uploads), T-2 and T-3, which is the class difference of a datacentre card behind a
|
|
327
|
+
cloud vCPU (host-side copies and submit latency), not a regression: the exact tier's `ms / iteration` is 1.7x the
|
|
328
|
+
RTX 4070 SUPER's at 10k and 3.0x at 65k.
|
|
329
|
+
|
|
330
|
+
Measured on gpu-linux-t4 (NVIDIA: 580.126.20 580.126.20.0), session 2026-09-18T07:03:17.146Z, medians of 5 runs; Chromium: nvidia / turing (nvidia-turing-driver0, the description is redacted by Chromium), session 2026-09-18T07:02:07.834Z.
|
|
331
|
+
|
|
332
|
+
| Id | What | Target | Measured |
|
|
333
|
+
| --- | -------------------------------------------------------------------------------------------- | ------------------- | --------------------- |
|
|
334
|
+
| T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 14.458 ms; 257.954 ms |
|
|
335
|
+
| T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 2.298 ms |
|
|
336
|
+
| T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 0.171 ms |
|
|
337
|
+
| T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.977 ms; 1.879 ms |
|
|
338
|
+
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.600 ms (1.357 ms) |
|
|
339
|
+
|
|
340
|
+
The exact curve (the `layout-exact` group: 2D, E = 10n, seeded G(n, m), one simulation per rung; ms / iteration from the profiler):
|
|
341
|
+
|
|
342
|
+
| n | ms / iteration | step(1) wall (ms) | pairs / s |
|
|
343
|
+
| ----- | -------------- | ----------------- | --------- |
|
|
344
|
+
| 1024 | 0.334 | 0.977 | 3.13e+9 |
|
|
345
|
+
| 4096 | 0.426 | 0.881 | 3.93e+10 |
|
|
346
|
+
| 8192 | 0.801 | 1.184 | 8.38e+10 |
|
|
347
|
+
| 10000 | 0.977 | 1.357 | 1.02e+11 |
|
|
348
|
+
| 16384 | 1.879 | 2.317 | 1.43e+11 |
|
|
349
|
+
| 32768 | 6.534 | 7.241 | 1.64e+11 |
|
|
350
|
+
| 65536 | 24.883 | 25.955 | 1.73e+11 |
|
|
351
|
+
|
|
318
352
|
## Development
|
|
319
353
|
|
|
320
354
|
```bash
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@graphty/webgpu-graph-algorithms",
|
|
3
|
-
"version": "0.2.
|
|
3
|
+
"version": "0.2.1",
|
|
4
4
|
"description": "WebGPU-accelerated graph algorithms and layouts over the @graphty/graph-format snapshot, for Node (Dawn) and browsers",
|
|
5
5
|
"author": "Adam Powers <apowers@ato.ms>",
|
|
6
6
|
"type": "module",
|
|
@@ -60,11 +60,11 @@
|
|
|
60
60
|
"homepage": "https://github.com/graphty-org/graphty-monorepo/tree/master/webgpu-graph-algorithms#readme",
|
|
61
61
|
"dependencies": {
|
|
62
62
|
"@webgpu/types": "^0.1.72",
|
|
63
|
-
"@graphty/graph-format": "^0.
|
|
63
|
+
"@graphty/graph-format": "^1.0.0"
|
|
64
64
|
},
|
|
65
65
|
"peerDependencies": {
|
|
66
66
|
"@graphty/algorithms": "^1.0.0",
|
|
67
|
-
"@graphty/graph-format": "^0.
|
|
67
|
+
"@graphty/graph-format": "^1.0.0",
|
|
68
68
|
"@graphty/layout": "^1.0.0",
|
|
69
69
|
"webgpu": ">=0.4.0 <1.0.0"
|
|
70
70
|
},
|