@graphty/webgpu-graph-algorithms 0.2.0 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +37 -3
  2. package/package.json +3 -3
package/README.md CHANGED
@@ -280,9 +280,12 @@ power of two; `benchmarks/layout-exact.bench.ts` `exactMaxNodesFromLadder`).
280
280
 
281
281
  ## Performance
282
282
 
283
- Regenerated from `benchmarks/results/nvidia-lovelace-driver580.json` (the last session) by the procedure
284
- recorded in `docs/decisions/G3.md` appendix A; the targets are the T-table of plan section 10.4. A missed target is re-fixed by a
285
- recorded owner decision in `docs/decisions/G<n>.md`, never relaxed silently.
283
+ Regenerated from the last session of each baseline under `benchmarks/results/` (`nvidia-lovelace-driver580.json`, the
284
+ dev box; `gpu-linux-t4.json`, the CI lane) by the procedure recorded in `docs/decisions/G3.md` appendix A; the targets
285
+ are the T-table of plan section 10.4. A missed target is re-fixed by a recorded owner decision in
286
+ `docs/decisions/G<n>.md`, never relaxed silently.
287
+
288
+ ### The dev box (nvidia-lovelace-driver580)
286
289
 
287
290
  Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-16T02:07:45.933Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z.
288
291
 
@@ -315,6 +318,37 @@ The end-to-end run of `benchmarks/layout-run.ts --nodes 100000 --edges 1000000`
315
318
  iterations, batches of 8) takes 18.971 ms per iteration on the same card, uploads and readbacks included (16.975 ms of
316
319
  GPU time per iteration in the last batch).
317
320
 
321
+ ### The CI lane (gpu-linux-t4)
322
+
323
+ The first run of the GPU lane (`gpu.yml`, graphty-monorepo run 35316416067, 2026-09-18) on a machine.dev T4 -- one Tesla
324
+ T4 (16 GB), 4 vCPU of a Xeon Platinum 8259CL, driver 580.126.20 -- wrote this baseline; `scripts/bench-compare.js` fails
325
+ a later run of the lane whose median exceeds 3x these figures. The T-table targets were set on the dev box; the T4 meets
326
+ T-4 and T-5 and misses T-1 (both uploads), T-2 and T-3, which is the class difference of a datacentre card behind a
327
+ cloud vCPU (host-side copies and submit latency), not a regression: the exact tier's `ms / iteration` is 1.7x the
328
+ RTX 4070 SUPER's at 10k and 3.0x at 65k.
329
+
330
+ Measured on gpu-linux-t4 (NVIDIA: 580.126.20 580.126.20.0), session 2026-09-18T07:03:17.146Z, medians of 5 runs; Chromium: nvidia / turing (nvidia-turing-driver0, the description is redacted by Chromium), session 2026-09-18T07:02:07.834Z.
331
+
332
+ | Id | What | Target | Measured |
333
+ | --- | -------------------------------------------------------------------------------------------- | ------------------- | --------------------- |
334
+ | T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 14.458 ms; 257.954 ms |
335
+ | T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 2.298 ms |
336
+ | T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 0.171 ms |
337
+ | T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.977 ms; 1.879 ms |
338
+ | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.600 ms (1.357 ms) |
339
+
340
+ The exact curve (the `layout-exact` group: 2D, E = 10n, seeded G(n, m), one simulation per rung; ms / iteration from the profiler):
341
+
342
+ | n | ms / iteration | step(1) wall (ms) | pairs / s |
343
+ | ----- | -------------- | ----------------- | --------- |
344
+ | 1024 | 0.334 | 0.977 | 3.13e+9 |
345
+ | 4096 | 0.426 | 0.881 | 3.93e+10 |
346
+ | 8192 | 0.801 | 1.184 | 8.38e+10 |
347
+ | 10000 | 0.977 | 1.357 | 1.02e+11 |
348
+ | 16384 | 1.879 | 2.317 | 1.43e+11 |
349
+ | 32768 | 6.534 | 7.241 | 1.64e+11 |
350
+ | 65536 | 24.883 | 25.955 | 1.73e+11 |
351
+
318
352
  ## Development
319
353
 
320
354
  ```bash
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@graphty/webgpu-graph-algorithms",
3
- "version": "0.2.0",
3
+ "version": "0.2.1",
4
4
  "description": "WebGPU-accelerated graph algorithms and layouts over the @graphty/graph-format snapshot, for Node (Dawn) and browsers",
5
5
  "author": "Adam Powers <apowers@ato.ms>",
6
6
  "type": "module",
@@ -60,11 +60,11 @@
60
60
  "homepage": "https://github.com/graphty-org/graphty-monorepo/tree/master/webgpu-graph-algorithms#readme",
61
61
  "dependencies": {
62
62
  "@webgpu/types": "^0.1.72",
63
- "@graphty/graph-format": "^0.2.1"
63
+ "@graphty/graph-format": "^1.0.0"
64
64
  },
65
65
  "peerDependencies": {
66
66
  "@graphty/algorithms": "^1.0.0",
67
- "@graphty/graph-format": "^0.2.0",
67
+ "@graphty/graph-format": "^1.0.0",
68
68
  "@graphty/layout": "^1.0.0",
69
69
  "webgpu": ">=0.4.0 <1.0.0"
70
70
  },