@graphty/webgpu-graph-algorithms 0.5.1 → 0.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +459 -58
- package/dist/browser.js +1 -1
- package/dist/chunks/{context-BR7fx3vR.js → context-BXqgCifx.js} +190 -40
- package/dist/chunks/context-BXqgCifx.js.map +1 -0
- package/dist/node.js +1 -1
- package/dist/src/algorithms/components.d.ts.map +1 -1
- package/dist/src/algorithms/components.js +12 -13
- package/dist/src/algorithms/components.js.map +1 -1
- package/dist/src/algorithms/degree.d.ts +6 -8
- package/dist/src/algorithms/degree.d.ts.map +1 -1
- package/dist/src/algorithms/degree.js +58 -35
- package/dist/src/algorithms/degree.js.map +1 -1
- package/dist/src/algorithms/pagerank.d.ts.map +1 -1
- package/dist/src/algorithms/pagerank.js +16 -14
- package/dist/src/algorithms/pagerank.js.map +1 -1
- package/dist/src/algorithms/power-iteration.d.ts +2 -2
- package/dist/src/algorithms/power-iteration.d.ts.map +1 -1
- package/dist/src/algorithms/power-iteration.js +17 -14
- package/dist/src/algorithms/power-iteration.js.map +1 -1
- package/dist/src/constants.d.ts +38 -8
- package/dist/src/constants.d.ts.map +1 -1
- package/dist/src/constants.js +38 -8
- package/dist/src/constants.js.map +1 -1
- package/dist/src/errors.d.ts +3 -2
- package/dist/src/errors.d.ts.map +1 -1
- package/dist/src/errors.js +2 -1
- package/dist/src/errors.js.map +1 -1
- package/dist/src/index.d.ts +6 -4
- package/dist/src/index.d.ts.map +1 -1
- package/dist/src/index.js +8 -3
- package/dist/src/index.js.map +1 -1
- package/dist/src/kernel/dispatch.d.ts +8 -3
- package/dist/src/kernel/dispatch.d.ts.map +1 -1
- package/dist/src/kernel/dispatch.js +18 -7
- package/dist/src/kernel/dispatch.js.map +1 -1
- package/dist/src/kernel/kernel.d.ts +30 -1
- package/dist/src/kernel/kernel.d.ts.map +1 -1
- package/dist/src/kernel/kernel.js +49 -5
- package/dist/src/kernel/kernel.js.map +1 -1
- package/dist/src/kernel/prelude.d.ts.map +1 -1
- package/dist/src/kernel/prelude.js +6 -1
- package/dist/src/kernel/prelude.js.map +1 -1
- package/dist/src/kernel/profiler.d.ts +15 -3
- package/dist/src/kernel/profiler.d.ts.map +1 -1
- package/dist/src/kernel/profiler.js +27 -4
- package/dist/src/kernel/profiler.js.map +1 -1
- package/dist/src/kernels.d.ts +17 -7
- package/dist/src/kernels.d.ts.map +1 -1
- package/dist/src/kernels.js +323 -16
- package/dist/src/kernels.js.map +1 -1
- package/dist/src/layouts/calibrate.d.ts +51 -0
- package/dist/src/layouts/calibrate.d.ts.map +1 -0
- package/dist/src/layouts/calibrate.js +172 -0
- package/dist/src/layouts/calibrate.js.map +1 -0
- package/dist/src/layouts/force-simulation.d.ts +39 -4
- package/dist/src/layouts/force-simulation.d.ts.map +1 -1
- package/dist/src/layouts/force-simulation.js +71 -19
- package/dist/src/layouts/force-simulation.js.map +1 -1
- package/dist/src/layouts/forceatlas2.d.ts +107 -36
- package/dist/src/layouts/forceatlas2.d.ts.map +1 -1
- package/dist/src/layouts/forceatlas2.js +296 -100
- package/dist/src/layouts/forceatlas2.js.map +1 -1
- package/dist/src/layouts/fruchterman-reingold.d.ts +73 -27
- package/dist/src/layouts/fruchterman-reingold.d.ts.map +1 -1
- package/dist/src/layouts/fruchterman-reingold.js +230 -70
- package/dist/src/layouts/fruchterman-reingold.js.map +1 -1
- package/dist/src/layouts/model-common.d.ts +41 -3
- package/dist/src/layouts/model-common.d.ts.map +1 -1
- package/dist/src/layouts/model-common.js +74 -3
- package/dist/src/layouts/model-common.js.map +1 -1
- package/dist/src/layouts/repulsion-grid.d.ts +152 -0
- package/dist/src/layouts/repulsion-grid.d.ts.map +1 -0
- package/dist/src/layouts/repulsion-grid.js +318 -0
- package/dist/src/layouts/repulsion-grid.js.map +1 -0
- package/dist/src/layouts/spring-electrical.d.ts +75 -30
- package/dist/src/layouts/spring-electrical.d.ts.map +1 -1
- package/dist/src/layouts/spring-electrical.js +231 -74
- package/dist/src/layouts/spring-electrical.js.map +1 -1
- package/dist/src/memory/residency.d.ts +6 -2
- package/dist/src/memory/residency.d.ts.map +1 -1
- package/dist/src/memory/residency.js +84 -14
- package/dist/src/memory/residency.js.map +1 -1
- package/dist/src/primitives/core-shape.d.ts +38 -2
- package/dist/src/primitives/core-shape.d.ts.map +1 -1
- package/dist/src/primitives/core-shape.js +71 -3
- package/dist/src/primitives/core-shape.js.map +1 -1
- package/dist/src/primitives/grid-pyramid.d.ts +71 -0
- package/dist/src/primitives/grid-pyramid.d.ts.map +1 -0
- package/dist/src/primitives/grid-pyramid.js +143 -0
- package/dist/src/primitives/grid-pyramid.js.map +1 -0
- package/dist/src/primitives/grid.d.ts +118 -0
- package/dist/src/primitives/grid.d.ts.map +1 -0
- package/dist/src/primitives/grid.js +225 -0
- package/dist/src/primitives/grid.js.map +1 -0
- package/dist/src/primitives/histogram.d.ts +67 -0
- package/dist/src/primitives/histogram.d.ts.map +1 -0
- package/dist/src/primitives/histogram.js +190 -0
- package/dist/src/primitives/histogram.js.map +1 -0
- package/dist/src/primitives/radix-sort.d.ts +75 -0
- package/dist/src/primitives/radix-sort.d.ts.map +1 -0
- package/dist/src/primitives/radix-sort.js +168 -0
- package/dist/src/primitives/radix-sort.js.map +1 -0
- package/dist/src/primitives/scan.d.ts +44 -0
- package/dist/src/primitives/scan.d.ts.map +1 -0
- package/dist/src/primitives/scan.js +151 -0
- package/dist/src/primitives/scan.js.map +1 -0
- package/dist/src/primitives/segmented-reduce.d.ts +25 -17
- package/dist/src/primitives/segmented-reduce.d.ts.map +1 -1
- package/dist/src/primitives/segmented-reduce.js +166 -47
- package/dist/src/primitives/segmented-reduce.js.map +1 -1
- package/dist/src/primitives/spmv.d.ts +18 -14
- package/dist/src/primitives/spmv.d.ts.map +1 -1
- package/dist/src/primitives/spmv.js +94 -58
- package/dist/src/primitives/spmv.js.map +1 -1
- package/dist/src/primitives/verify.d.ts +49 -0
- package/dist/src/primitives/verify.d.ts.map +1 -0
- package/dist/src/primitives/verify.js +229 -0
- package/dist/src/primitives/verify.js.map +1 -0
- package/dist/src/types/context.d.ts +53 -0
- package/dist/src/types/context.d.ts.map +1 -1
- package/dist/src/types/layout.d.ts +20 -0
- package/dist/src/types/layout.d.ts.map +1 -1
- package/dist/src/wgsl/counting-scatter.wgsl.d.ts +8 -0
- package/dist/src/wgsl/counting-scatter.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/counting-scatter.wgsl.js +17 -0
- package/dist/src/wgsl/counting-scatter.wgsl.js.map +1 -0
- package/dist/src/wgsl/fa2-attraction.wgsl.d.ts +23 -11
- package/dist/src/wgsl/fa2-attraction.wgsl.d.ts.map +1 -1
- package/dist/src/wgsl/fa2-attraction.wgsl.js +98 -20
- package/dist/src/wgsl/fa2-attraction.wgsl.js.map +1 -1
- package/dist/src/wgsl/fa2-stats-finalize.wgsl.d.ts +6 -2
- package/dist/src/wgsl/fa2-stats-finalize.wgsl.d.ts.map +1 -1
- package/dist/src/wgsl/fa2-stats-finalize.wgsl.js +22 -1
- package/dist/src/wgsl/fa2-stats-finalize.wgsl.js.map +1 -1
- package/dist/src/wgsl/grid-cell-key.wgsl.d.ts +8 -0
- package/dist/src/wgsl/grid-cell-key.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-cell-key.wgsl.js +30 -0
- package/dist/src/wgsl/grid-cell-key.wgsl.js.map +1 -0
- package/dist/src/wgsl/grid-centroid-hub.wgsl.d.ts +8 -0
- package/dist/src/wgsl/grid-centroid-hub.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-centroid-hub.wgsl.js +29 -0
- package/dist/src/wgsl/grid-centroid-hub.wgsl.js.map +1 -0
- package/dist/src/wgsl/grid-centroid.wgsl.d.ts +8 -0
- package/dist/src/wgsl/grid-centroid.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-centroid.wgsl.js +29 -0
- package/dist/src/wgsl/grid-centroid.wgsl.js.map +1 -0
- package/dist/src/wgsl/grid-downsample.wgsl.d.ts +7 -0
- package/dist/src/wgsl/grid-downsample.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-downsample.wgsl.js +28 -0
- package/dist/src/wgsl/grid-downsample.wgsl.js.map +1 -0
- package/dist/src/wgsl/grid-far-field.wgsl.d.ts +13 -0
- package/dist/src/wgsl/grid-far-field.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-far-field.wgsl.js +98 -0
- package/dist/src/wgsl/grid-far-field.wgsl.js.map +1 -0
- package/dist/src/wgsl/grid-near-field.wgsl.d.ts +19 -0
- package/dist/src/wgsl/grid-near-field.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/grid-near-field.wgsl.js +129 -0
- package/dist/src/wgsl/grid-near-field.wgsl.js.map +1 -0
- package/dist/src/wgsl/histogram.wgsl.d.ts +7 -0
- package/dist/src/wgsl/histogram.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/histogram.wgsl.js +15 -0
- package/dist/src/wgsl/histogram.wgsl.js.map +1 -0
- package/dist/src/wgsl/indirect-finalize.wgsl.d.ts +8 -0
- package/dist/src/wgsl/indirect-finalize.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/indirect-finalize.wgsl.js +26 -0
- package/dist/src/wgsl/indirect-finalize.wgsl.js.map +1 -0
- package/dist/src/wgsl/radix-hist.wgsl.d.ts +9 -0
- package/dist/src/wgsl/radix-hist.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/radix-hist.wgsl.js +31 -0
- package/dist/src/wgsl/radix-hist.wgsl.js.map +1 -0
- package/dist/src/wgsl/radix-scatter.wgsl.d.ts +9 -0
- package/dist/src/wgsl/radix-scatter.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/radix-scatter.wgsl.js +40 -0
- package/dist/src/wgsl/radix-scatter.wgsl.js.map +1 -0
- package/dist/src/wgsl/scan-add.wgsl.d.ts +6 -0
- package/dist/src/wgsl/scan-add.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/scan-add.wgsl.js +14 -0
- package/dist/src/wgsl/scan-add.wgsl.js.map +1 -0
- package/dist/src/wgsl/scan-block.wgsl.d.ts +8 -0
- package/dist/src/wgsl/scan-block.wgsl.d.ts.map +1 -0
- package/dist/src/wgsl/scan-block.wgsl.js +30 -0
- package/dist/src/wgsl/scan-block.wgsl.js.map +1 -0
- package/dist/src/wgsl/segmented-reduce.wgsl.d.ts +22 -8
- package/dist/src/wgsl/segmented-reduce.wgsl.d.ts.map +1 -1
- package/dist/src/wgsl/segmented-reduce.wgsl.js +84 -15
- package/dist/src/wgsl/segmented-reduce.wgsl.js.map +1 -1
- package/dist/src/wgsl/spmv-pull.wgsl.d.ts +22 -11
- package/dist/src/wgsl/spmv-pull.wgsl.d.ts.map +1 -1
- package/dist/src/wgsl/spmv-pull.wgsl.js +110 -36
- package/dist/src/wgsl/spmv-pull.wgsl.js.map +1 -1
- package/dist/tsconfig.build.tsbuildinfo +1 -1
- package/dist/webgpu-graph-algorithms.js +3815 -1003
- package/dist/webgpu-graph-algorithms.js.map +1 -1
- package/package.json +9 -8
- package/src/algorithms/components.ts +12 -16
- package/src/algorithms/degree.ts +58 -43
- package/src/algorithms/pagerank.ts +20 -18
- package/src/algorithms/power-iteration.ts +19 -18
- package/src/constants.ts +38 -8
- package/src/errors.ts +3 -1
- package/src/index.ts +14 -4
- package/src/kernel/dispatch.ts +18 -7
- package/src/kernel/kernel.ts +59 -5
- package/src/kernel/prelude.ts +9 -0
- package/src/kernel/profiler.ts +28 -4
- package/src/kernels.ts +356 -18
- package/src/layouts/calibrate.ts +187 -0
- package/src/layouts/force-simulation.ts +91 -23
- package/src/layouts/forceatlas2.ts +331 -106
- package/src/layouts/fruchterman-reingold.ts +255 -74
- package/src/layouts/model-common.ts +98 -3
- package/src/layouts/repulsion-grid.ts +451 -0
- package/src/layouts/spring-electrical.ts +257 -78
- package/src/memory/residency.ts +126 -20
- package/src/primitives/core-shape.ts +91 -4
- package/src/primitives/grid-pyramid.ts +221 -0
- package/src/primitives/grid.ts +349 -0
- package/src/primitives/histogram.ts +273 -0
- package/src/primitives/radix-sort.ts +246 -0
- package/src/primitives/scan.ts +197 -0
- package/src/primitives/segmented-reduce.ts +214 -56
- package/src/primitives/spmv.ts +125 -65
- package/src/primitives/verify.ts +249 -0
- package/src/types/context.ts +56 -0
- package/src/types/layout.ts +22 -0
- package/src/wgsl/counting-scatter.wgsl.ts +16 -0
- package/src/wgsl/fa2-attraction.wgsl.ts +98 -20
- package/src/wgsl/fa2-stats-finalize.wgsl.ts +22 -1
- package/src/wgsl/grid-cell-key.wgsl.ts +29 -0
- package/src/wgsl/grid-centroid-hub.wgsl.ts +28 -0
- package/src/wgsl/grid-centroid.wgsl.ts +28 -0
- package/src/wgsl/grid-downsample.wgsl.ts +27 -0
- package/src/wgsl/grid-far-field.wgsl.ts +97 -0
- package/src/wgsl/grid-near-field.wgsl.ts +128 -0
- package/src/wgsl/histogram.wgsl.ts +14 -0
- package/src/wgsl/indirect-finalize.wgsl.ts +25 -0
- package/src/wgsl/radix-hist.wgsl.ts +30 -0
- package/src/wgsl/radix-scatter.wgsl.ts +39 -0
- package/src/wgsl/scan-add.wgsl.ts +13 -0
- package/src/wgsl/scan-block.wgsl.ts +29 -0
- package/src/wgsl/segmented-reduce.wgsl.ts +84 -15
- package/src/wgsl/spmv-pull.wgsl.ts +110 -36
- package/dist/chunks/context-BR7fx3vR.js.map +0 -1
package/README.md
CHANGED
|
@@ -3,20 +3,24 @@
|
|
|
3
3
|
WebGPU-accelerated graph algorithms and layouts over the `@graphty/graph-format` snapshot, for Node
|
|
4
4
|
(Dawn, through the `webgpu` npm package) and browsers (Chromium). One code base, three entry points:
|
|
5
5
|
|
|
6
|
-
| Entry | Import | What it gives you
|
|
7
|
-
| ------------------------------------------ | ------------- |
|
|
8
|
-
| `@graphty/webgpu-graph-algorithms` | the core | `createForceAtlas2`, `
|
|
9
|
-
| `@graphty/webgpu-graph-algorithms/node` | Node only | `createNodeGpuContext`, `probeNodeWebGpu`, `createNodeGpu` (Dawn), `dawnFlags`
|
|
10
|
-
| `@graphty/webgpu-graph-algorithms/browser` | browsers only | `probeBrowserWebGpu`, `requestGpuContext`
|
|
11
|
-
|
|
12
|
-
**Status:
|
|
13
|
-
|
|
14
|
-
`
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
6
|
+
| Entry | Import | What it gives you |
|
|
7
|
+
| ------------------------------------------ | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
8
|
+
| `@graphty/webgpu-graph-algorithms` | the core | the layouts (`createForceAtlas2`, `createFruchtermanReingold`, `createSpringElectrical`, `seedPositions`), the algorithms (`pageRank`, `personalizedPageRank`, `hits`, `eigenvectorCentrality`, `katzCentrality`, `connectedComponents`, `degree`), `createAccelerator`, `calibrateLayout`, `verifyDevice`, `GpuContext`, `WebGpuGraphError`, `isSoftwareAdapter`, the constants (`EXACT_MAX_NODES`, `FA2_DEFAULTS`, `FR_DEFAULTS`, `SE_DEFAULTS`, `LAYOUT_TUNING_DEFAULTS`, ...) and the option / stats / accelerator types |
|
|
9
|
+
| `@graphty/webgpu-graph-algorithms/node` | Node only | `createNodeGpuContext`, `probeNodeWebGpu`, `createNodeGpu` (Dawn), `dawnFlags` |
|
|
10
|
+
| `@graphty/webgpu-graph-algorithms/browser` | browsers only | `probeBrowserWebGpu`, `requestGpuContext` |
|
|
11
|
+
|
|
12
|
+
**Status: three force layouts and six algorithms, on the exact and the grid repulsion tiers.**
|
|
13
|
+
ForceAtlas2, Fruchterman-Reingold and ngraph's spring-electrical preset run from Node (`run()`) and from a
|
|
14
|
+
browser frame loop (`step()` once per frame) at any size: the default `repulsion: "auto"` runs the exact tier
|
|
15
|
+
up to `exactMaxNodes` = 32768 nodes (owner decision G4-D1 in `docs/decisions/G4.md`) and the grid tier above
|
|
16
|
+
it; `repulsion: "exact"` forces all pairs at any size (O(n^2) per iteration: 18.971 ms per iteration at 100k
|
|
17
|
+
nodes / 1M edges on an RTX 4070 SUPER) and `repulsion: "grid"` forces the grid tier. PageRank, personalized
|
|
18
|
+
PageRank, HITS, eigenvector centrality, Katz centrality and weakly connected components run as plain async
|
|
19
|
+
calls on the same context. The phases follow the phase plan of
|
|
20
|
+
`design/webgpu/webgpu-acceleration-plan.md` (monorepo root) section 13 and the interface contract
|
|
21
|
+
`design/webgpu/plans/2026-09-14-webgpu-p0-p3-interfaces.md`; the gate records are `docs/decisions/G<n>.md`.
|
|
22
|
+
There is no CPU fallback anywhere in this package: when no adapter or device exists it throws
|
|
23
|
+
`WebGpuGraphError`.
|
|
20
24
|
|
|
21
25
|
## Install
|
|
22
26
|
|
|
@@ -61,8 +65,8 @@ ctx.release(snapshot); // destroys the snapshot's buffers (nothing is freed by G
|
|
|
61
65
|
ctx.dispose(); // destroys the device and lets the process exit
|
|
62
66
|
```
|
|
63
67
|
|
|
64
|
-
|
|
65
|
-
|
|
68
|
+
Above `exactMaxNodes` the default `"auto"` picks the grid tier (spec 7.8: by n alone); `repulsion: "exact"`
|
|
69
|
+
keeps the all-pairs tier at any size and `repulsion: "grid"` picks the grid tier at any size. The same run
|
|
66
70
|
from the command line, with a verification of the result:
|
|
67
71
|
|
|
68
72
|
```bash
|
|
@@ -85,7 +89,7 @@ if (!probe.ok) {
|
|
|
85
89
|
throw new Error(`${probe.code}: ${probe.reason ?? ""}`); // E_NO_WEBGPU, E_NO_ADAPTER or E_SOFTWARE_ONLY
|
|
86
90
|
}
|
|
87
91
|
const ctx = await requestGpuContext({ adapter: probe.adapter ?? undefined });
|
|
88
|
-
const acc = createAccelerator(ctx, { layout: { exactMaxNodes:
|
|
92
|
+
const acc = createAccelerator(ctx, { layout: { exactMaxNodes: 4096 } }); // the tuning every simulation inherits
|
|
89
93
|
const sim = acc.forceAtlas2({ seed: 1, iterationsPerStep: 1, maxInFlight: 2 });
|
|
90
94
|
sim.load(snapshot, positions); // positions: your stride-3 Float32Array; NaN rows are seeded
|
|
91
95
|
|
|
@@ -146,13 +150,13 @@ names and defaults as the CPU port in `@graphty/layout`) plus the GPU tuning:
|
|
|
146
150
|
GPU tuning (`GpuLayoutTuning`; also the `layout` field of `createAccelerator`'s options, inherited by every
|
|
147
151
|
simulation the accelerator creates):
|
|
148
152
|
|
|
149
|
-
| Option | Default | Meaning
|
|
150
|
-
| --------------------------------------------------- | ----------------------- |
|
|
151
|
-
| `repulsion` | `"auto"` | `"exact"` at any n; `"auto"` = exact iff n <= `exactMaxNodes
|
|
152
|
-
| `exactMaxNodes` | `32768` | the
|
|
153
|
-
| `deterministic` | `true` | fixed summation order (the exact tier is always deterministic)
|
|
154
|
-
| `compat` | `"paper"` | `"networkx"` reproduces NetworkX 3.4's `forceatlas2_layout` (gravity toward the origin, its accumulated swing / traction)
|
|
155
|
-
| `nearMax`, `gridMax2D`, `gridMax3D`, `extentFactor` | `64`, `512`, `128`, `6` | stored for the grid tier (P4)
|
|
153
|
+
| Option | Default | Meaning |
|
|
154
|
+
| --------------------------------------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
155
|
+
| `repulsion` | `"auto"` | `"exact"` at any n; `"auto"` = exact iff n <= `exactMaxNodes`, else grid; `"grid"` at any n |
|
|
156
|
+
| `exactMaxNodes` | `32768` | the G3 value, kept by owner decision G4-D1 while the accuracy work G4-F1 leaves is open (`docs/decisions/G4.md`; the section 7.8 rule computes 1024 on the RTX 4070 SUPER); pass your own (`calibrateLayout` measures it). |
|
|
157
|
+
| `deterministic` | `true` | fixed summation order (the exact tier is always deterministic) |
|
|
158
|
+
| `compat` | `"paper"` | `"networkx"` reproduces NetworkX 3.4's `forceatlas2_layout` (gravity toward the origin, its accumulated swing / traction) |
|
|
159
|
+
| `nearMax`, `gridMax2D`, `gridMax3D`, `extentFactor` | `64`, `512`, `128`, `6` | stored for the grid tier (P4) |
|
|
156
160
|
|
|
157
161
|
Every range error is `E_INVALID_ARGUMENT` (`gravity < 0`, `scalingRatio <= 0`, `jitterTolerance <= 0`,
|
|
158
162
|
`maxIter < 1`, `settleWindow < 1`, `maxInFlight < 1`, `iterationsPerStep < 1`, `dim` not 2 or 3,
|
|
@@ -169,6 +173,150 @@ else the batch's wall time divided by its iteration count), the grid fields (`nu
|
|
|
169
173
|
`trace`: one `{ swing, traction, speed, speedEfficiency, meanDisplacement, settledCount }` record per
|
|
170
174
|
iteration of the last batch.
|
|
171
175
|
|
|
176
|
+
## Fruchterman-Reingold and the spring-electrical preset
|
|
177
|
+
|
|
178
|
+
Two more force layouts ship beside ForceAtlas2, on the same exact repulsion tier and with exactly the
|
|
179
|
+
lifecycle the two ForceAtlas2 sections above describe: `load(snapshot, positions)`, then `run()` from Node or
|
|
180
|
+
`step()` once per frame, with `stats`, `settled`, `setPosition`, `setFixed`, `reheat`, `flush`, `setParams`,
|
|
181
|
+
`dispose` and the coalescing rule all unchanged. They differ in the force law, in their option records and in
|
|
182
|
+
what their trace carries. Both are also methods of an accelerator: `acc.fruchtermanReingold(options)` and
|
|
183
|
+
`acc.springElectrical(options)`, which pass the accelerator's `layout` tuning down the same way
|
|
184
|
+
`acc.forceAtlas2` does.
|
|
185
|
+
|
|
186
|
+
What they ask of the graph, and what they refuse, is ForceAtlas2's:
|
|
187
|
+
|
|
188
|
+
- an UNDIRECTED snapshot. `load()` throws `E_SNAPSHOT { reason: "directed" }` otherwise; pass
|
|
189
|
+
`toUndirected().snapshot`
|
|
190
|
+
- edge weights are ignored by both models, and both give every node the same mass rule for its whole run --
|
|
191
|
+
1 for Fruchterman-Reingold, `1 + degree / 3` for the spring preset -- so neither has a `weight` or
|
|
192
|
+
`nodeMass` option to pass
|
|
193
|
+
- `positions` is the owner's stride-3 scene-unit `Float32Array` of length `3 * nodeCount`, over a plain
|
|
194
|
+
`ArrayBuffer`; any other length is `E_INVALID_ARGUMENT`. Rows holding a non-finite component are seeded in
|
|
195
|
+
place at `load()` (see `seedPositions` below), finite rows are the starting layout
|
|
196
|
+
- an empty snapshot (`nodeCount: 0`) loads without touching the GPU and is `settled` on arrival; `step()`
|
|
197
|
+
resolves immediately and the positions array is left alone
|
|
198
|
+
- a graph the device cannot hold is refused at `load()`, before any GPU work:
|
|
199
|
+
`E_TOO_LARGE { path: "positions" }` when `16 * nodeCount` bytes exceed the device's `maxBufferSize`,
|
|
200
|
+
`{ path: "windowed" }` when the arc arrays need more than one storage binding, and `{ path: "partials" }`
|
|
201
|
+
above 16,776,960 nodes
|
|
202
|
+
- above `exactMaxNodes` with the default `repulsion: "auto"`, `load()` throws
|
|
203
|
+
`E_UNSUPPORTED { feature: "repulsion.grid" }`, as ForceAtlas2 does; `repulsion: "exact"` runs all pairs at
|
|
204
|
+
any size
|
|
205
|
+
|
|
206
|
+
The GPU tuning table above applies to both, with one exception: `compat` is read by ForceAtlas2 alone and
|
|
207
|
+
changes nothing here.
|
|
208
|
+
|
|
209
|
+
### Fruchterman-Reingold
|
|
210
|
+
|
|
211
|
+
The spring model of the original paper: every pair of nodes repels with `k^2 / d`, every edge pulls with
|
|
212
|
+
`d^2 / k`, where `k` is the ideal edge length, and a falling temperature caps how far a node may move in a
|
|
213
|
+
single iteration.
|
|
214
|
+
|
|
215
|
+
```ts
|
|
216
|
+
import { createFruchtermanReingold } from "@graphty/webgpu-graph-algorithms";
|
|
217
|
+
|
|
218
|
+
// positions: the owner's stride-3 scene-unit array; NaN rows are seeded at load(), finite rows are kept
|
|
219
|
+
const positions = new Float32Array(3 * snapshot.nodeCount).fill(Number.NaN);
|
|
220
|
+
|
|
221
|
+
const sim = createFruchtermanReingold(ctx, { iterations: 50, cooling: "linear", seed: 42, dim: 2 });
|
|
222
|
+
sim.load(snapshot, positions); // an undirected snapshot
|
|
223
|
+
const stats = await sim.run({ batch: 8 }); // stops at `iterations` or when the layout settles
|
|
224
|
+
console.log(sim.iterationsDone, sim.settled, stats.temperature, stats.meanDisplacement);
|
|
225
|
+
|
|
226
|
+
sim.dispose(); // then ctx.release(snapshot) and ctx.dispose() when the graph goes away
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
| Option | Default | Meaning |
|
|
230
|
+
| --------------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
231
|
+
| `k` | `null` | the ideal edge length. `null` (and `0`, and `NaN`) means `1 / sqrt(nodeCount)`, resolved at `load()`; anything else must be a finite number > 0 |
|
|
232
|
+
| `iterations` | `50` | an integer >= 0 and the budget `run()` stops at (`0` is settled on arrival). Under `cooling: "adaptive"` it is only a cap, and an option record that never set it gets `10000` instead |
|
|
233
|
+
| `cooling` | `"linear"` | `"linear"`: the temperature falls from `0.1` to `0` across `iterations`, so a run lasts the whole budget. `"adaptive"`: Yifan Hu's step control -- the temperature shrinks by `0.9` whenever the total force energy rose and grows by `1 / 0.9` after five consecutive falls -- so a run ends when the layout stops moving rather than at a fixed count |
|
|
234
|
+
| `fixed` | `null` | which nodes are pinned, applied at `load()`: a graph-format `NodeMask`, the name of a bool node column, or `null`, which takes the column with role `fixed` when the snapshot has one and pins nothing otherwise. Not a live option: `setParams({ fixed })` is `E_INVALID_ARGUMENT`, use `setFixed(mask)` |
|
|
235
|
+
| `dim`, `scale`, `center`, `seed` | `2`, `1`, `[0, 0, 0]`, `null` | as ForceAtlas2 |
|
|
236
|
+
| `settleThreshold`, `settleWindow`, `iterationsPerStep`, `maxInFlight` | `0.001`, `10`, `1`, `2` | as ForceAtlas2 |
|
|
237
|
+
|
|
238
|
+
`FR_DEFAULTS` is the frozen record of those defaults (`center` and `seed` are the shared ones and are not in
|
|
239
|
+
it). Every range error is `E_INVALID_ARGUMENT`, and `maxInFlight` cannot change after creation.
|
|
240
|
+
|
|
241
|
+
`sim.stats` is `FruchtermanReingoldStats`: everything in `LayoutStatsBase` -- `iteration`,
|
|
242
|
+
`meanDisplacement`, `rmsRadius`, `layoutRadius`, `centroid`, `repulsionTier` (`"exact"`), `msPerIteration`
|
|
243
|
+
and the grid fields, which are `null` -- plus `temperature`, the schedule's value for the last completed
|
|
244
|
+
iteration, and `trace`: one `{ temperature, meanDisplacement, settledCount }` record per iteration of the
|
|
245
|
+
last batch.
|
|
246
|
+
|
|
247
|
+
### The spring-electrical preset
|
|
248
|
+
|
|
249
|
+
ngraph.forcelayout's model with ngraph's own constants: Hooke springs along the edges, Coulomb repulsion
|
|
250
|
+
between every pair, and a velocity integrated with drag. A node's mass is `1 + degree / 3`, as ngraph
|
|
251
|
+
assigns it.
|
|
252
|
+
|
|
253
|
+
```ts
|
|
254
|
+
import { createSpringElectrical } from "@graphty/webgpu-graph-algorithms";
|
|
255
|
+
|
|
256
|
+
const sim = createSpringElectrical(ctx, { springLength: 10, timeStep: 0.5, seed: 42 });
|
|
257
|
+
sim.load(snapshot, positions);
|
|
258
|
+
const stats = await sim.run({ maxIter: 1000, batch: 8 }); // pass maxIter: this model has no iteration count
|
|
259
|
+
console.log(sim.iterationsDone, sim.settled, stats.kineticEnergy);
|
|
260
|
+
|
|
261
|
+
sim.dispose();
|
|
262
|
+
```
|
|
263
|
+
|
|
264
|
+
| Option | Default | Meaning |
|
|
265
|
+
| --------------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
266
|
+
| `springLength` | `10` | the rest length of an edge; a finite number > 0 |
|
|
267
|
+
| `springCoefficient` | `0.8`, size-scaled | Hooke's constant. Left out or `null` it is ngraph's `0.8` times `min(1, 300 / nodeCount)`, applied once `nodeCount` is known at `load()`; a number you pass is used as given at every size |
|
|
268
|
+
| `gravity` | `-12`, size-scaled | ngraph's Coulomb constant, where NEGATIVE repels -- this is not ForceAtlas2's centre gravity. Left out or `null` it is ngraph's `-12` times the same factor. Any finite number is accepted |
|
|
269
|
+
| `dragCoefficient` | `0.9` | velocity damping; a finite number >= 0 |
|
|
270
|
+
| `timeStep` | `0.5` | the integrator's step; a finite number > 0 |
|
|
271
|
+
| `dim`, `scale`, `center`, `seed` | `2`, `1`, `[0, 0, 0]`, `null` | as ForceAtlas2 |
|
|
272
|
+
| `settleThreshold`, `settleWindow`, `iterationsPerStep`, `maxInFlight` | `0.001`, `10`, `1`, `2` | as ForceAtlas2 |
|
|
273
|
+
|
|
274
|
+
The size scaling exists because ngraph's constants were tuned for graphs of a few hundred nodes: on tens of
|
|
275
|
+
thousands the per-node forces are large enough that every node moves at the unit speed clamp and the layout
|
|
276
|
+
never comes to rest. `SE_DEFAULTS` is the frozen record of the unscaled constants.
|
|
277
|
+
|
|
278
|
+
This model has no iteration-count option, so `run()` without `maxIter` has no budget at all and returns only
|
|
279
|
+
when the layout settles. Pass `run({ maxIter })` unless that is what you want.
|
|
280
|
+
|
|
281
|
+
`sim.stats` is `SpringElectricalStats`: `LayoutStatsBase` plus `kineticEnergy`, and a `trace` of
|
|
282
|
+
`{ kineticEnergy, meanDisplacement, settledCount }`. The energy lags the positions by one iteration -- the
|
|
283
|
+
iteration that follows an integrate is the one that folds its energy -- so the first record after `load()`
|
|
284
|
+
carries `0` and record i carries the energy of iteration i - 1.
|
|
285
|
+
|
|
286
|
+
Both models appear in the Performance section below as row T-14 and in the `layout-fr` benchmark group.
|
|
287
|
+
|
|
288
|
+
## The `seedPositions` helper
|
|
289
|
+
|
|
290
|
+
Every simulation seeds the unplaced rows of your positions array at `load()`, so you never have to call this.
|
|
291
|
+
It is public for the case where you want the same starting layout without a simulation -- to draw the graph
|
|
292
|
+
before the first frame lands, to reproduce a run on the CPU, or to place a subgraph the same way twice.
|
|
293
|
+
|
|
294
|
+
```ts
|
|
295
|
+
import { fromEdgeArrays } from "@graphty/graph-format";
|
|
296
|
+
import { seedPositions } from "@graphty/webgpu-graph-algorithms";
|
|
297
|
+
|
|
298
|
+
const s = fromEdgeArrays({ directed: false, nodeCount: 3, src: new Uint32Array([0]), dst: new Uint32Array([1]) });
|
|
299
|
+
const positions = new Float32Array(3 * s.nodeCount).fill(Number.NaN);
|
|
300
|
+
|
|
301
|
+
seedPositions(s, positions, 42, 2, 1, null, "fa2");
|
|
302
|
+
// every row now holds x and y in [-1, 1) and z = 0; seed 42 gives this same array every time
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
The arguments are positional and all required: the snapshot, the owner's stride-3 scene-unit array (modified
|
|
306
|
+
in place, length `3 * nodeCount`), the LCG seed (`0` or `null` draws a random one, the CPU port's quirk kept
|
|
307
|
+
bit for bit), `dim` (`2` or `3`), the scene `scale` (> 0), the scene `center` (an `ArrayLike<number>`, or
|
|
308
|
+
`null` for the origin), and the draw range: `"fa2"` is `[-1, 1)` in layout units, `"fr"` is `[0, 1)`.
|
|
309
|
+
ForceAtlas2 and the spring preset seed with `"fa2"`, Fruchterman-Reingold with `"fr"`.
|
|
310
|
+
|
|
311
|
+
A row counts as unseeded when ANY of its first `dim` components is not finite, and only its non-finite
|
|
312
|
+
components are written -- a finite component is never changed and a fully finite row is never touched. In 2D
|
|
313
|
+
the third component of a row being seeded is set to `center[2]`. The draw box depends on what is already
|
|
314
|
+
there: when no row is fully finite the draw is the plain range per axis, and otherwise it is the `[min, max]`
|
|
315
|
+
box of the finite components, per axis, so new nodes land among the ones already placed instead of around the
|
|
316
|
+
origin. No random number is drawn at all when nothing needs seeding. A bad `dim`, a `positions` length other
|
|
317
|
+
than `3 * nodeCount`, a non-positive `scale`
|
|
318
|
+
or a non-finite `center` component is `E_INVALID_ARGUMENT`; the function returns nothing.
|
|
319
|
+
|
|
172
320
|
## Acquisition
|
|
173
321
|
|
|
174
322
|
### Node
|
|
@@ -243,10 +391,218 @@ ctx.dispose(); // destroys the device and lets the process exit
|
|
|
243
391
|
In a browser, `import { requestGpuContext } from "@graphty/webgpu-graph-algorithms/browser"` and `const ctx = await
|
|
244
392
|
requestGpuContext();` replace the first import and line; the rest is identical.
|
|
245
393
|
|
|
394
|
+
## Centrality and components
|
|
395
|
+
|
|
396
|
+
Six algorithm functions run on the device. Every one has the same shape -- `await fn(ctx, snapshot, options?)`
|
|
397
|
+
-- and every array that comes back is indexed by the snapshot's NODE INDEX (`0 .. nodeCount - 1`), never by
|
|
398
|
+
node id; the snapshot's id map turns an index back into the id it was built from. Every score result carries
|
|
399
|
+
`precision: "f32"`, so a reader can label what it is looking at.
|
|
400
|
+
|
|
401
|
+
The context uploads what an algorithm needs the first time it sees a snapshot and keeps it until
|
|
402
|
+
`ctx.release(snapshot)`: the CSR core for PageRank, HITS, eigenvector centrality and connected components,
|
|
403
|
+
the reverse adjacency for PageRank, HITS and Katz, the edge list for connected components. A second call on
|
|
404
|
+
the same snapshot pays no upload.
|
|
405
|
+
|
|
406
|
+
Each one is also a method of the object `createAccelerator(ctx)` returns -- `acc.pageRank(s, options)` and so
|
|
407
|
+
on, with `connectedComponents` answering to `weaklyConnectedComponents` as well -- which is how
|
|
408
|
+
`@graphty/algorithms` reaches them when a GPU accelerator is injected into it.
|
|
409
|
+
|
|
410
|
+
The PageRank example below is complete; the examples after it reuse its context `ctx` and its snapshot `s`
|
|
411
|
+
rather than repeating the setup.
|
|
412
|
+
|
|
413
|
+
Options every one of them accepts on top of its own (`GpuRunOptions`):
|
|
414
|
+
|
|
415
|
+
| Option | Default | Meaning |
|
|
416
|
+
| ------------ | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
417
|
+
| `dest` | unset | a preallocated result of exactly `nodeCount` elements over a plain `ArrayBuffer`: a `Float32Array` for scores, a `Uint32Array` for labels. Any other type or length is `E_INVALID_ARGUMENT { argument: "dest" }`; `hits` fills it with the hubs |
|
|
418
|
+
| `signal` | unset | an `AbortSignal`, checked before the first submit and after every batch; `E_ABORTED { batchId }` |
|
|
419
|
+
| `onProgress` | unset | `(done, total)` after every batch. `total` is `maxIterations`, `2 * maxIterations` for `hits` (it runs two chains) and `1` for `connectedComponents`, which reports once when it finishes |
|
|
420
|
+
|
|
421
|
+
A graph too large for the device is refused before any GPU work: when the arc arrays need more than one
|
|
422
|
+
storage binding they would need a windowed upload, which only `degree` executes, and each of these six throws
|
|
423
|
+
`E_TOO_LARGE { needed, limit, path: "windowed" }` instead. A node count whose dispatch does not fit
|
|
424
|
+
the device's 2D workgroup grid is `E_TOO_LARGE { path: "dispatch" }`, and scratch that does not fit the
|
|
425
|
+
device's budget is `E_OUT_OF_MEMORY { requested, resident, label }`.
|
|
426
|
+
|
|
427
|
+
### PageRank
|
|
428
|
+
|
|
429
|
+
PageRank is the stationary distribution of a surfer who follows an out-arc with probability `dampingFactor`
|
|
430
|
+
and teleports to a uniformly random node otherwise. `pageRank` runs NetworkX's iteration on the device: rank
|
|
431
|
+
is pulled along the reverse adjacency, out-weights normalise the push, and the mass sitting on nodes with no
|
|
432
|
+
out-arcs is redistributed through the teleport vector every iteration.
|
|
433
|
+
|
|
434
|
+
```ts
|
|
435
|
+
import { fromEdgeArrays } from "@graphty/graph-format";
|
|
436
|
+
import { pageRank } from "@graphty/webgpu-graph-algorithms";
|
|
437
|
+
import { createNodeGpuContext } from "@graphty/webgpu-graph-algorithms/node";
|
|
438
|
+
|
|
439
|
+
const ctx = await createNodeGpuContext();
|
|
440
|
+
const s = fromEdgeArrays({
|
|
441
|
+
directed: true,
|
|
442
|
+
nodeCount: 4,
|
|
443
|
+
src: new Uint32Array([0, 1, 2, 3]), // a tail into a 3-cycle: 0 -> 1 -> 2 -> 3 -> 1
|
|
444
|
+
dst: new Uint32Array([1, 2, 3, 1]),
|
|
445
|
+
});
|
|
446
|
+
|
|
447
|
+
const r = await pageRank(ctx, s, { dampingFactor: 0.85, tolerance: 1e-6 });
|
|
448
|
+
console.log(r.scores); // Float32Array(4), index-aligned, summing to 1; node 0 scores lowest
|
|
449
|
+
console.log(r.iterations, r.converged, r.danglingMass);
|
|
450
|
+
|
|
451
|
+
ctx.release(s); // when the graph goes away; ctx.dispose() at the end of the program
|
|
452
|
+
```
|
|
453
|
+
|
|
454
|
+
| Option | Default | Meaning |
|
|
455
|
+
| --------------- | ------- | -------------------------------------------------------------------------------------------------------------------- |
|
|
456
|
+
| `dampingFactor` | `0.85` | the probability of following an arc; the rest teleports |
|
|
457
|
+
| `maxIterations` | `100` | an integer >= 1; anything else is `E_INVALID_ARGUMENT` |
|
|
458
|
+
| `tolerance` | `1e-6` | converged when the L1 change of the whole vector falls below `tolerance * nodeCount` (NetworkX's rule) |
|
|
459
|
+
| `weighted` | `true` | arc weights are the transition mass; `false` weighs every arc 1. A snapshot with no weights is unweighted either way |
|
|
460
|
+
|
|
461
|
+
`GpuPageRankResult` comes back: `scores` (one f32 per node, summing to 1 up to f32 rounding), `iterations`
|
|
462
|
+
(the FIRST iteration whose delta fell below the threshold, not the batch boundary the run stopped at, and
|
|
463
|
+
`maxIterations` when it never did), `converged`, `danglingMass` (the rank mass the last iteration found on
|
|
464
|
+
nodes with no positive out-weight and redistributed) and `precision`.
|
|
465
|
+
|
|
466
|
+
Directed or undirected both work: an undirected snapshot carries both directions, so the pull and the
|
|
467
|
+
normaliser see the same arcs, and the scores differ from the directed form's as they should. `nodeCount: 0`
|
|
468
|
+
returns an empty `scores` with `iterations: 0`, `converged: true` and does no GPU work. A graph with no arcs
|
|
469
|
+
at all makes every node dangling, so the teleport vector is already the fixed point: the call returns
|
|
470
|
+
`1 / nodeCount` everywhere with `iterations: 0` and `danglingMass: 1`. An isolated node inside a larger graph
|
|
471
|
+
is one dangling node -- it keeps its teleport share and its share of the redistributed mass -- and so is a
|
|
472
|
+
node whose out-arcs all weigh zero under `weighted: true`.
|
|
473
|
+
|
|
474
|
+
### Personalized PageRank
|
|
475
|
+
|
|
476
|
+
The same iteration with the uniform teleport vector replaced by yours, so the walk restarts where you say:
|
|
477
|
+
|
|
478
|
+
```ts
|
|
479
|
+
import { personalizedPageRank } from "@graphty/webgpu-graph-algorithms";
|
|
480
|
+
|
|
481
|
+
const bias = new Float32Array(s.nodeCount);
|
|
482
|
+
bias[0] = 1; // restart at node 0 only
|
|
483
|
+
const r = await personalizedPageRank(ctx, s, bias, { dampingFactor: 0.85 });
|
|
484
|
+
console.log(r.scores); // mass concentrated on what node 0 reaches
|
|
485
|
+
```
|
|
486
|
+
|
|
487
|
+
`personalization` is a `Float32Array` of one finite non-negative number per node, not all zero, and it is
|
|
488
|
+
normalised to sum 1 on the host before the run -- so unnormalised weights are fine. A wrong length, a
|
|
489
|
+
negative or non-finite entry, or a zero total is `E_INVALID_ARGUMENT { argument: "personalization" }`.
|
|
490
|
+
Options, result and edge cases are PageRank's, except that a graph with no arcs returns the normalised
|
|
491
|
+
personalization vector rather than `1 / nodeCount`.
|
|
492
|
+
|
|
493
|
+
### HITS
|
|
494
|
+
|
|
495
|
+
HITS scores each node twice: as a hub (it points at good authorities) and as an authority (good hubs point at
|
|
496
|
+
it). `hits` runs the CPU package's recurrence -- `a(i) = A^T norm(h(i-1))` and `h(i) = A norm(a(i-1))` from
|
|
497
|
+
uniform seeds -- as two interleaved chains on the device, then sum-normalises both vectors once on the host.
|
|
498
|
+
|
|
499
|
+
```ts
|
|
500
|
+
import { hits } from "@graphty/webgpu-graph-algorithms";
|
|
501
|
+
|
|
502
|
+
const cycle = fromEdgeArrays({
|
|
503
|
+
directed: true,
|
|
504
|
+
nodeCount: 2,
|
|
505
|
+
src: new Uint32Array([0, 1]),
|
|
506
|
+
dst: new Uint32Array([1, 0]),
|
|
507
|
+
});
|
|
508
|
+
const r = await hits(ctx, cycle);
|
|
509
|
+
console.log(r.hubs, r.authorities); // [0.5, 0.5] and [0.5, 0.5]
|
|
510
|
+
ctx.release(cycle);
|
|
511
|
+
```
|
|
512
|
+
|
|
513
|
+
The options are `maxIterations` (`100`), `tolerance` (`1e-6`) and `weighted` (`true`), with PageRank's
|
|
514
|
+
meanings. `GpuHitsResult` carries `hubs` and `authorities` (both f32, index-aligned, each summing to 1),
|
|
515
|
+
`iterations` (the larger of the two chains'), `converged` (both chains) and `precision`. `dest` receives the
|
|
516
|
+
hubs; the authorities always come back in a fresh array.
|
|
517
|
+
|
|
518
|
+
Direction is the point of the algorithm: an undirected snapshot has a symmetric adjacency, so the two vectors
|
|
519
|
+
come out the same. `nodeCount: 0` returns two empty arrays with `iterations: 0` and `converged: true`; a graph
|
|
520
|
+
with no arcs returns all zeros in both, since there is nothing to be a hub of.
|
|
521
|
+
|
|
522
|
+
### Eigenvector centrality
|
|
523
|
+
|
|
524
|
+
A node is central when central nodes point at it. The scores are the principal eigenvector of the adjacency
|
|
525
|
+
matrix, found by power iteration over the forward adjacency and L2-normalised once at the end.
|
|
526
|
+
|
|
527
|
+
```ts
|
|
528
|
+
import { eigenvectorCentrality } from "@graphty/webgpu-graph-algorithms";
|
|
529
|
+
|
|
530
|
+
const r = await eigenvectorCentrality(ctx, s, { maxIterations: 100, tolerance: 1e-6 });
|
|
531
|
+
console.log(r.scores, r.converged); // f32, index-aligned, L2 norm 1
|
|
532
|
+
```
|
|
533
|
+
|
|
534
|
+
The options are `maxIterations` (`100`), `tolerance` (`1e-6`) and `weighted` (`true`); convergence is
|
|
535
|
+
PageRank's L1 rule, a delta below `tolerance * nodeCount`. It returns `GpuScoresResult` -- `scores`,
|
|
536
|
+
`iterations`, `converged`, `precision` -- which is PageRank's result without `danglingMass`.
|
|
537
|
+
|
|
538
|
+
Directed or undirected both work. When a graph's components have different spectral radii the power iteration
|
|
539
|
+
converges on the dominant component's eigenvector and the others fall toward zero; that is a property of the
|
|
540
|
+
method, not of this implementation. `nodeCount: 0` is empty; a graph with no arcs, and an isolated node in a
|
|
541
|
+
larger graph, scores 0.
|
|
542
|
+
|
|
543
|
+
### Katz centrality
|
|
544
|
+
|
|
545
|
+
Katz counts every walk that ends at a node, discounted by its length: `x = alpha * A^T x + beta`, so a walk of
|
|
546
|
+
length L contributes `alpha^L`. Unlike eigenvector centrality it hands every node the floor `beta`, which
|
|
547
|
+
keeps a node with no incoming arcs from scoring zero.
|
|
548
|
+
|
|
549
|
+
```ts
|
|
550
|
+
import { katzCentrality } from "@graphty/webgpu-graph-algorithms";
|
|
551
|
+
|
|
552
|
+
const r = await katzCentrality(ctx, s, { alpha: 0.1, beta: 1 });
|
|
553
|
+
console.log(r.scores); // f32, index-aligned, L2-normalised
|
|
554
|
+
```
|
|
555
|
+
|
|
556
|
+
| Option | Default | Meaning |
|
|
557
|
+
| --------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
|
|
558
|
+
| `alpha` | `0.1` | the attenuation per step. Convergence needs it below the reciprocal of the largest eigenvalue; the package checks only that it is finite |
|
|
559
|
+
| `beta` | `1` | the constant every node is given each iteration |
|
|
560
|
+
| `maxIterations` | `100` | as above |
|
|
561
|
+
| `tolerance` | `1e-6` | as above |
|
|
562
|
+
| `weighted` | `true` | as above |
|
|
563
|
+
|
|
564
|
+
A non-finite `alpha` or `beta` is `E_INVALID_ARGUMENT`. The result is `GpuScoresResult`, L2-normalised on the
|
|
565
|
+
host with no per-iteration normaliser on the device. Walks arrive along the reverse adjacency, so direction
|
|
566
|
+
matters. `nodeCount: 0` is empty; with no arcs every node holds `beta` alone, which after normalisation is
|
|
567
|
+
`1 / sqrt(nodeCount)` each.
|
|
568
|
+
|
|
569
|
+
### Connected components
|
|
570
|
+
|
|
571
|
+
`connectedComponents` labels the WEAK components: it walks the edge list, so an edge joins its two endpoints
|
|
572
|
+
whether or not the snapshot is directed, and the directed and undirected forms of one edge set give identical
|
|
573
|
+
labels. The kernel is GAP's Afforest -- sampled link rounds, then each edge once, then a final compress.
|
|
574
|
+
|
|
575
|
+
```ts
|
|
576
|
+
import { connectedComponents } from "@graphty/webgpu-graph-algorithms";
|
|
577
|
+
|
|
578
|
+
const g = fromEdgeArrays({
|
|
579
|
+
directed: false,
|
|
580
|
+
nodeCount: 4,
|
|
581
|
+
src: new Uint32Array([0, 1]), // the path 0 - 1 - 2, with node 3 on its own
|
|
582
|
+
dst: new Uint32Array([1, 2]),
|
|
583
|
+
});
|
|
584
|
+
const r = await connectedComponents(ctx, g);
|
|
585
|
+
console.log(r.labels); // Uint32Array [0, 0, 0, 1]
|
|
586
|
+
console.log(r.count); // 2
|
|
587
|
+
console.log(r.groups()); // [Uint32Array [0, 1, 2], Uint32Array [3]]
|
|
588
|
+
ctx.release(g);
|
|
589
|
+
```
|
|
590
|
+
|
|
591
|
+
`labels` holds one label per node index. By default they are renumbered dense `0 .. count - 1` in first-seen
|
|
592
|
+
index order, which makes them identical to the labels `@graphty/algorithms` produces for the same graph;
|
|
593
|
+
`renumber: false` returns the raw root indices instead -- the same partition with arbitrary label values, and
|
|
594
|
+
the same `count`. `groups()` builds the member lists on the first call and caches them, index-aligned with
|
|
595
|
+
the labels in that same first-seen order.
|
|
596
|
+
|
|
597
|
+
`renumber` (default `true`) is the only option besides the shared three. `nodeCount: 0` gives empty labels,
|
|
598
|
+
`count: 0` and no GPU work; a graph with no arcs gives every node its own block, so `labels[v] === v` and
|
|
599
|
+
`count === nodeCount`, which is what an isolated node gets inside a larger graph too. Self-loops and repeated
|
|
600
|
+
edges change nothing: the link step is idempotent.
|
|
601
|
+
|
|
246
602
|
## Errors
|
|
247
603
|
|
|
248
604
|
Every condition the package detects itself is a `WebGpuGraphError` with a stable `code`
|
|
249
|
-
(`E_NO_WEBGPU`, `E_NO_ADAPTER`, `E_NO_DEVICE`, `E_SOFTWARE_ONLY`, `E_DEVICE_LOST`, `E_DISPOSED`,
|
|
605
|
+
(`E_NO_WEBGPU`, `E_NO_ADAPTER`, `E_NO_DEVICE`, `E_SOFTWARE_ONLY`, `E_DEVICE_LOST`, `E_DEVICE_INCORRECT`, `E_DISPOSED`,
|
|
250
606
|
`E_VALIDATION`, `E_SHADER_COMPILE`, `E_OUT_OF_MEMORY`, `E_TOO_LARGE`, `E_UNSUPPORTED`,
|
|
251
607
|
`E_INVALID_ARGUMENT`, `E_SNAPSHOT`, `E_RELEASED`, `E_NOT_LOADED`, `E_ABORTED`) and frozen `details`.
|
|
252
608
|
`isWebGpuGraphError(x)` and `hasErrorCode(x, code)` are structural brand checks, so they survive two copies
|
|
@@ -255,6 +611,38 @@ of the package. The graph-format codes `E_GPU_INELIGIBLE`, `E_UNKNOWN_NODE`, `E_
|
|
|
255
611
|
released rejects its next `step()` with `E_RELEASED`; a lost device disposes every simulation and rejects
|
|
256
612
|
every pending `step()` with `E_DEVICE_LOST` (create a new context from a fresh adapter and `load()` again).
|
|
257
613
|
|
|
614
|
+
## The device self-check
|
|
615
|
+
|
|
616
|
+
Some GPU drivers return wrong answers rather than failing. On Windows over the Microsoft Basic Render Driver,
|
|
617
|
+
compute shaders that synchronise across a workgroup produce silently incorrect results, which would make every
|
|
618
|
+
number this package computes there unreliable with no error anywhere.
|
|
619
|
+
|
|
620
|
+
So the first algorithm or layout you run on a context asks the device for an answer this package already knows --
|
|
621
|
+
an exclusive scan across 33 workgroups of known numbers (8,193 words on a 256-lane device), verified word by word
|
|
622
|
+
on the host -- and refuses a device that gets it wrong with `E_DEVICE_INCORRECT`, naming the first wrong word,
|
|
623
|
+
what belonged there and the adapter's description string. It runs once per device and is remembered: 14-20 ms
|
|
624
|
+
the first time, nearly all of it compiling the two scan pipelines that any scan-using algorithm would compile
|
|
625
|
+
anyway, and 0.6-1.0 ms of work under that (measured on Dawn over lavapipe and over an RTX 4070 SUPER). It cannot be switched off: a flag for it would be off in
|
|
626
|
+
somebody's production build, which is the silent wrong answers walking back in.
|
|
627
|
+
|
|
628
|
+
Choosing the processor instead is your decision, not the package's, so nothing falls back. To ask before you
|
|
629
|
+
commit work to a device, call it yourself:
|
|
630
|
+
|
|
631
|
+
```ts
|
|
632
|
+
import { verifyDevice } from "@graphty/webgpu-graph-algorithms";
|
|
633
|
+
|
|
634
|
+
const check = await verifyDevice(ctx); // memoised: the algorithms below reuse this result
|
|
635
|
+
if (!check.ok) {
|
|
636
|
+
// check.mismatch names the first wrong word; check.description is the adapter string that identifies the driver
|
|
637
|
+
runOnTheCpuInstead();
|
|
638
|
+
}
|
|
639
|
+
```
|
|
640
|
+
|
|
641
|
+
It checks one property -- that values crossing a workgroup barrier, and block totals crossing dispatches of one
|
|
642
|
+
compute pass, survive. A device that gets that right and gets atomics or float rounding wrong still passes. It is
|
|
643
|
+
a refusal mechanism, not a certificate of correctness; `docs/decisions/device-self-check.md` records what it
|
|
644
|
+
covers, what it does not, and why the package refuses such a device rather than computing around it.
|
|
645
|
+
|
|
258
646
|
## Benchmarks
|
|
259
647
|
|
|
260
648
|
```bash
|
|
@@ -277,10 +665,15 @@ the profiler reports; every rung starts with an untimed clock warm-up burst, bec
|
|
|
277
665
|
SM clock at its idle 210 MHz under sparse sub-millisecond dispatches and the kernels then measure 4-15x slower),
|
|
278
666
|
`pagerank` (T-8), `wcc` (T-9), `layout-fr` (T-14: `step(1)` of `createFruchtermanReingold` and `createSpringElectrical`
|
|
279
667
|
at 10k and at 100k with `repulsion: "exact"`, the same two rows per model and rung as `layout-exact`, tagged `fr` /
|
|
280
|
-
`se`)
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
668
|
+
`se`), `layout-grid` (T-6 and T-7: `step(1)` of the grid tier on the grid ladder 32k / 65k / 100k / 262k / 1M in 2D
|
|
669
|
+
and in 3D, the same two rows per rung tagged `grid` with the dimension, plus the `fa2-attraction` pass of the 1M 2D
|
|
670
|
+
iteration from the profiler, and the exact ladder's 1k / 4k / 8k / 16k rungs in 2D for the crossover re-check). The Chromium numbers of T-5 (10k on the exact tier, 100k on the grid tier) come from the
|
|
671
|
+
`bench`-tagged browser test (`GRAPHTY_BROWSER_GPU=nvidia node scripts/run-browser-project.js`), which appends its
|
|
672
|
+
session through the Vitest commands bridge. `exactMaxNodes` is re-fixed from the ladder by the rule of plan section 7.8
|
|
673
|
+
(the largest rung under 4 ms per iteration and not slower than the grid tier at the same n -- the `layout-grid` 2D rows,
|
|
674
|
+
which the group also records at the exact ladder's 1k / 4k / 8k / 16k rungs so the clause has a row at every rung --
|
|
675
|
+
rounded down to a power of two; `benchmarks/layout-exact.bench.ts` `exactMaxNodesFromLadder`); `calibrateLayout(ctx)`
|
|
676
|
+
measures the same crossover on a consumer's own device.
|
|
284
677
|
|
|
285
678
|
## Performance
|
|
286
679
|
|
|
@@ -291,18 +684,21 @@ are the T-table of plan section 10.4. A missed target is re-fixed by a recorded
|
|
|
291
684
|
|
|
292
685
|
### The dev box (nvidia-lovelace-driver580)
|
|
293
686
|
|
|
294
|
-
Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-20T01:06:48.656Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z.
|
|
295
|
-
|
|
296
|
-
| Id | What | Target
|
|
297
|
-
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------- |
|
|
298
|
-
| T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms
|
|
299
|
-
| T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms
|
|
300
|
-
| T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms
|
|
301
|
-
| T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms
|
|
302
|
-
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms
|
|
303
|
-
| T-
|
|
304
|
-
| T-
|
|
305
|
-
| T-
|
|
687
|
+
Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-20T01:06:48.656Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z. The T-5 (100k), T-6 and T-7 rows are from the later session 2026-09-21T06:15:17.027Z (the file's last session, the current baseline, the one the crossover re-check reads; its Chromium session 2026-09-21T05:48:14.190Z), whose other rows are within 1.25x of this table's.
|
|
688
|
+
|
|
689
|
+
| Id | What | Target | Measured |
|
|
690
|
+
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- | ---------------------------- |
|
|
691
|
+
| T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 5.980 ms; 125.738 ms |
|
|
692
|
+
| T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 0.870 ms |
|
|
693
|
+
| T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 0.181 ms |
|
|
694
|
+
| T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.586 ms; 1.053 ms |
|
|
695
|
+
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.400 ms (0.726 ms) |
|
|
696
|
+
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 100k on the grid tier, Chromium (Node in brackets) | <= 12 ms | 3.700 ms (2.087 ms) |
|
|
697
|
+
| T-6 | ForceAtlas2 grid tier, GPU time per iteration (profiler) at 100k 2D; at 1M 2D; at 100k 3D | <= 10 ms; <= 100 ms; <= 20 ms | 0.634 ms; 5.414 ms; 1.306 ms |
|
|
698
|
+
| T-7 | Attraction gather (the `fa2-attraction` pass of the grid tier), GPU time per iteration (profiler) at 1M / 10M | <= 15 ms | 1.493 ms |
|
|
699
|
+
| T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 17.204 ms; 198.645 ms |
|
|
700
|
+
| T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 145.970 ms (12.613 ms) |
|
|
701
|
+
| T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.617 ms; 16.367 ms |
|
|
306
702
|
|
|
307
703
|
Three rows miss their target in this session: the 1M / 10M upload (125.7 ms against 100 ms, the open owner decision of
|
|
308
704
|
`docs/decisions/G1.md` section 7), the empty-submit round trip (0.181 ms against 0.1 ms: the row is measured after the
|
|
@@ -336,23 +732,28 @@ GPU time per iteration in the last batch).
|
|
|
336
732
|
|
|
337
733
|
The first run of the GPU lane (`gpu.yml`, graphty-monorepo run 35316416067, 2026-09-18) on a machine.dev T4 -- one Tesla
|
|
338
734
|
T4 (16 GB), 4 vCPU of a Xeon Platinum 8259CL, driver 580.126.20 -- wrote this baseline; `scripts/bench-compare.js` fails
|
|
339
|
-
a later run of the lane whose median
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
|
349
|
-
|
|
|
350
|
-
| T-
|
|
351
|
-
| T-
|
|
352
|
-
| T-
|
|
353
|
-
| T-
|
|
354
|
-
| T-
|
|
355
|
-
| T-
|
|
735
|
+
a later run of the lane whose median AND minimum both exceed 1.35x the best figures this file has ever held, by at
|
|
736
|
+
least 2.5 ms. The T-table targets were set on the dev box; the T4 meets
|
|
737
|
+
T-4, T-5 and T-6 and misses T-1 (both uploads), T-2, T-3 and T-7 (the 1M attraction pass of the grid tier, 18.165 ms
|
|
738
|
+
against 15 ms), which is the class difference of a datacentre card behind a cloud vCPU (host-side copies and submit
|
|
739
|
+
latency), not a regression: the exact tier's `ms / iteration` is 1.7x the RTX 4070 SUPER's at 10k and 3.0x at 65k, and
|
|
740
|
+
the grid tier's is 2.8x at 100k 2D and 5.7x at 1M 2D while the attraction pass alone is 12.2x.
|
|
741
|
+
|
|
742
|
+
Measured on gpu-linux-t4 (NVIDIA: 580.126.20 580.126.20.0), session 2026-09-20T02:29:33.210Z (run 35483512705), medians of 5 runs; Chromium: nvidia / turing (nvidia-turing-driver0, the description is redacted by Chromium), session 2026-09-20T02:28:10.676Z. The T-5 (100k), T-6 and T-7 rows are from the later session 2026-09-22T20:21:53.814Z (GPU lane run 35775999450 on commit 1d2d4dcf, the file's last session and the current baseline of this class; its Chromium session 2026-09-22T20:19:30.117Z).
|
|
743
|
+
|
|
744
|
+
| Id | What | Target | Measured |
|
|
745
|
+
| ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------- |
|
|
746
|
+
| T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 14.338 ms; 256.418 ms |
|
|
747
|
+
| T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 2.957 ms |
|
|
748
|
+
| T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 1.296 ms |
|
|
749
|
+
| T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.972 ms; 1.953 ms |
|
|
750
|
+
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.600 ms (1.359 ms) |
|
|
751
|
+
| T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 100k on the grid tier, Chromium (Node in brackets) | <= 12 ms | 8.100 ms (5.314 ms) |
|
|
752
|
+
| T-6 | ForceAtlas2 grid tier, GPU time per iteration (profiler) at 100k 2D; at 1M 2D; at 100k 3D | <= 10 ms; <= 100 ms; <= 20 ms | 1.790 ms; 31.057 ms; 3.954 ms |
|
|
753
|
+
| T-7 | Attraction gather (the `fa2-attraction` pass of the grid tier), GPU time per iteration (profiler) at 1M / 10M | <= 15 ms | 18.165 ms -- MISSED: 21 % over the target on this card (1.493 ms on the dev box; G4-F16 of docs/decisions/G4.md) |
|
|
754
|
+
| T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 45.461 ms; 1092.799 ms |
|
|
755
|
+
| T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 293.079 ms (28.544 ms) |
|
|
756
|
+
| T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.942 ms; 54.232 ms (the spring preset 1.051 ms; 60.706 ms) |
|
|
356
757
|
|
|
357
758
|
The exact curve (the `layout-exact` group: 2D, E = 10n, seeded G(n, m), one simulation per rung; ms / iteration from the profiler):
|
|
358
759
|
|