@graphty/webgpu-graph-algorithms 0.5.1 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (243) hide show
  1. package/README.md +459 -58
  2. package/dist/browser.js +1 -1
  3. package/dist/chunks/{context-BR7fx3vR.js → context-BXqgCifx.js} +190 -40
  4. package/dist/chunks/context-BXqgCifx.js.map +1 -0
  5. package/dist/node.js +1 -1
  6. package/dist/src/algorithms/components.d.ts.map +1 -1
  7. package/dist/src/algorithms/components.js +12 -13
  8. package/dist/src/algorithms/components.js.map +1 -1
  9. package/dist/src/algorithms/degree.d.ts +6 -8
  10. package/dist/src/algorithms/degree.d.ts.map +1 -1
  11. package/dist/src/algorithms/degree.js +58 -35
  12. package/dist/src/algorithms/degree.js.map +1 -1
  13. package/dist/src/algorithms/pagerank.d.ts.map +1 -1
  14. package/dist/src/algorithms/pagerank.js +16 -14
  15. package/dist/src/algorithms/pagerank.js.map +1 -1
  16. package/dist/src/algorithms/power-iteration.d.ts +2 -2
  17. package/dist/src/algorithms/power-iteration.d.ts.map +1 -1
  18. package/dist/src/algorithms/power-iteration.js +17 -14
  19. package/dist/src/algorithms/power-iteration.js.map +1 -1
  20. package/dist/src/constants.d.ts +38 -8
  21. package/dist/src/constants.d.ts.map +1 -1
  22. package/dist/src/constants.js +38 -8
  23. package/dist/src/constants.js.map +1 -1
  24. package/dist/src/errors.d.ts +3 -2
  25. package/dist/src/errors.d.ts.map +1 -1
  26. package/dist/src/errors.js +2 -1
  27. package/dist/src/errors.js.map +1 -1
  28. package/dist/src/index.d.ts +6 -4
  29. package/dist/src/index.d.ts.map +1 -1
  30. package/dist/src/index.js +8 -3
  31. package/dist/src/index.js.map +1 -1
  32. package/dist/src/kernel/dispatch.d.ts +8 -3
  33. package/dist/src/kernel/dispatch.d.ts.map +1 -1
  34. package/dist/src/kernel/dispatch.js +18 -7
  35. package/dist/src/kernel/dispatch.js.map +1 -1
  36. package/dist/src/kernel/kernel.d.ts +30 -1
  37. package/dist/src/kernel/kernel.d.ts.map +1 -1
  38. package/dist/src/kernel/kernel.js +49 -5
  39. package/dist/src/kernel/kernel.js.map +1 -1
  40. package/dist/src/kernel/prelude.d.ts.map +1 -1
  41. package/dist/src/kernel/prelude.js +6 -1
  42. package/dist/src/kernel/prelude.js.map +1 -1
  43. package/dist/src/kernel/profiler.d.ts +15 -3
  44. package/dist/src/kernel/profiler.d.ts.map +1 -1
  45. package/dist/src/kernel/profiler.js +27 -4
  46. package/dist/src/kernel/profiler.js.map +1 -1
  47. package/dist/src/kernels.d.ts +17 -7
  48. package/dist/src/kernels.d.ts.map +1 -1
  49. package/dist/src/kernels.js +323 -16
  50. package/dist/src/kernels.js.map +1 -1
  51. package/dist/src/layouts/calibrate.d.ts +51 -0
  52. package/dist/src/layouts/calibrate.d.ts.map +1 -0
  53. package/dist/src/layouts/calibrate.js +172 -0
  54. package/dist/src/layouts/calibrate.js.map +1 -0
  55. package/dist/src/layouts/force-simulation.d.ts +39 -4
  56. package/dist/src/layouts/force-simulation.d.ts.map +1 -1
  57. package/dist/src/layouts/force-simulation.js +71 -19
  58. package/dist/src/layouts/force-simulation.js.map +1 -1
  59. package/dist/src/layouts/forceatlas2.d.ts +107 -36
  60. package/dist/src/layouts/forceatlas2.d.ts.map +1 -1
  61. package/dist/src/layouts/forceatlas2.js +296 -100
  62. package/dist/src/layouts/forceatlas2.js.map +1 -1
  63. package/dist/src/layouts/fruchterman-reingold.d.ts +73 -27
  64. package/dist/src/layouts/fruchterman-reingold.d.ts.map +1 -1
  65. package/dist/src/layouts/fruchterman-reingold.js +230 -70
  66. package/dist/src/layouts/fruchterman-reingold.js.map +1 -1
  67. package/dist/src/layouts/model-common.d.ts +41 -3
  68. package/dist/src/layouts/model-common.d.ts.map +1 -1
  69. package/dist/src/layouts/model-common.js +74 -3
  70. package/dist/src/layouts/model-common.js.map +1 -1
  71. package/dist/src/layouts/repulsion-grid.d.ts +152 -0
  72. package/dist/src/layouts/repulsion-grid.d.ts.map +1 -0
  73. package/dist/src/layouts/repulsion-grid.js +318 -0
  74. package/dist/src/layouts/repulsion-grid.js.map +1 -0
  75. package/dist/src/layouts/spring-electrical.d.ts +75 -30
  76. package/dist/src/layouts/spring-electrical.d.ts.map +1 -1
  77. package/dist/src/layouts/spring-electrical.js +231 -74
  78. package/dist/src/layouts/spring-electrical.js.map +1 -1
  79. package/dist/src/memory/residency.d.ts +6 -2
  80. package/dist/src/memory/residency.d.ts.map +1 -1
  81. package/dist/src/memory/residency.js +84 -14
  82. package/dist/src/memory/residency.js.map +1 -1
  83. package/dist/src/primitives/core-shape.d.ts +38 -2
  84. package/dist/src/primitives/core-shape.d.ts.map +1 -1
  85. package/dist/src/primitives/core-shape.js +71 -3
  86. package/dist/src/primitives/core-shape.js.map +1 -1
  87. package/dist/src/primitives/grid-pyramid.d.ts +71 -0
  88. package/dist/src/primitives/grid-pyramid.d.ts.map +1 -0
  89. package/dist/src/primitives/grid-pyramid.js +143 -0
  90. package/dist/src/primitives/grid-pyramid.js.map +1 -0
  91. package/dist/src/primitives/grid.d.ts +118 -0
  92. package/dist/src/primitives/grid.d.ts.map +1 -0
  93. package/dist/src/primitives/grid.js +225 -0
  94. package/dist/src/primitives/grid.js.map +1 -0
  95. package/dist/src/primitives/histogram.d.ts +67 -0
  96. package/dist/src/primitives/histogram.d.ts.map +1 -0
  97. package/dist/src/primitives/histogram.js +190 -0
  98. package/dist/src/primitives/histogram.js.map +1 -0
  99. package/dist/src/primitives/radix-sort.d.ts +75 -0
  100. package/dist/src/primitives/radix-sort.d.ts.map +1 -0
  101. package/dist/src/primitives/radix-sort.js +168 -0
  102. package/dist/src/primitives/radix-sort.js.map +1 -0
  103. package/dist/src/primitives/scan.d.ts +44 -0
  104. package/dist/src/primitives/scan.d.ts.map +1 -0
  105. package/dist/src/primitives/scan.js +151 -0
  106. package/dist/src/primitives/scan.js.map +1 -0
  107. package/dist/src/primitives/segmented-reduce.d.ts +25 -17
  108. package/dist/src/primitives/segmented-reduce.d.ts.map +1 -1
  109. package/dist/src/primitives/segmented-reduce.js +166 -47
  110. package/dist/src/primitives/segmented-reduce.js.map +1 -1
  111. package/dist/src/primitives/spmv.d.ts +18 -14
  112. package/dist/src/primitives/spmv.d.ts.map +1 -1
  113. package/dist/src/primitives/spmv.js +94 -58
  114. package/dist/src/primitives/spmv.js.map +1 -1
  115. package/dist/src/primitives/verify.d.ts +49 -0
  116. package/dist/src/primitives/verify.d.ts.map +1 -0
  117. package/dist/src/primitives/verify.js +229 -0
  118. package/dist/src/primitives/verify.js.map +1 -0
  119. package/dist/src/types/context.d.ts +53 -0
  120. package/dist/src/types/context.d.ts.map +1 -1
  121. package/dist/src/types/layout.d.ts +20 -0
  122. package/dist/src/types/layout.d.ts.map +1 -1
  123. package/dist/src/wgsl/counting-scatter.wgsl.d.ts +8 -0
  124. package/dist/src/wgsl/counting-scatter.wgsl.d.ts.map +1 -0
  125. package/dist/src/wgsl/counting-scatter.wgsl.js +17 -0
  126. package/dist/src/wgsl/counting-scatter.wgsl.js.map +1 -0
  127. package/dist/src/wgsl/fa2-attraction.wgsl.d.ts +23 -11
  128. package/dist/src/wgsl/fa2-attraction.wgsl.d.ts.map +1 -1
  129. package/dist/src/wgsl/fa2-attraction.wgsl.js +98 -20
  130. package/dist/src/wgsl/fa2-attraction.wgsl.js.map +1 -1
  131. package/dist/src/wgsl/fa2-stats-finalize.wgsl.d.ts +6 -2
  132. package/dist/src/wgsl/fa2-stats-finalize.wgsl.d.ts.map +1 -1
  133. package/dist/src/wgsl/fa2-stats-finalize.wgsl.js +22 -1
  134. package/dist/src/wgsl/fa2-stats-finalize.wgsl.js.map +1 -1
  135. package/dist/src/wgsl/grid-cell-key.wgsl.d.ts +8 -0
  136. package/dist/src/wgsl/grid-cell-key.wgsl.d.ts.map +1 -0
  137. package/dist/src/wgsl/grid-cell-key.wgsl.js +30 -0
  138. package/dist/src/wgsl/grid-cell-key.wgsl.js.map +1 -0
  139. package/dist/src/wgsl/grid-centroid-hub.wgsl.d.ts +8 -0
  140. package/dist/src/wgsl/grid-centroid-hub.wgsl.d.ts.map +1 -0
  141. package/dist/src/wgsl/grid-centroid-hub.wgsl.js +29 -0
  142. package/dist/src/wgsl/grid-centroid-hub.wgsl.js.map +1 -0
  143. package/dist/src/wgsl/grid-centroid.wgsl.d.ts +8 -0
  144. package/dist/src/wgsl/grid-centroid.wgsl.d.ts.map +1 -0
  145. package/dist/src/wgsl/grid-centroid.wgsl.js +29 -0
  146. package/dist/src/wgsl/grid-centroid.wgsl.js.map +1 -0
  147. package/dist/src/wgsl/grid-downsample.wgsl.d.ts +7 -0
  148. package/dist/src/wgsl/grid-downsample.wgsl.d.ts.map +1 -0
  149. package/dist/src/wgsl/grid-downsample.wgsl.js +28 -0
  150. package/dist/src/wgsl/grid-downsample.wgsl.js.map +1 -0
  151. package/dist/src/wgsl/grid-far-field.wgsl.d.ts +13 -0
  152. package/dist/src/wgsl/grid-far-field.wgsl.d.ts.map +1 -0
  153. package/dist/src/wgsl/grid-far-field.wgsl.js +98 -0
  154. package/dist/src/wgsl/grid-far-field.wgsl.js.map +1 -0
  155. package/dist/src/wgsl/grid-near-field.wgsl.d.ts +19 -0
  156. package/dist/src/wgsl/grid-near-field.wgsl.d.ts.map +1 -0
  157. package/dist/src/wgsl/grid-near-field.wgsl.js +129 -0
  158. package/dist/src/wgsl/grid-near-field.wgsl.js.map +1 -0
  159. package/dist/src/wgsl/histogram.wgsl.d.ts +7 -0
  160. package/dist/src/wgsl/histogram.wgsl.d.ts.map +1 -0
  161. package/dist/src/wgsl/histogram.wgsl.js +15 -0
  162. package/dist/src/wgsl/histogram.wgsl.js.map +1 -0
  163. package/dist/src/wgsl/indirect-finalize.wgsl.d.ts +8 -0
  164. package/dist/src/wgsl/indirect-finalize.wgsl.d.ts.map +1 -0
  165. package/dist/src/wgsl/indirect-finalize.wgsl.js +26 -0
  166. package/dist/src/wgsl/indirect-finalize.wgsl.js.map +1 -0
  167. package/dist/src/wgsl/radix-hist.wgsl.d.ts +9 -0
  168. package/dist/src/wgsl/radix-hist.wgsl.d.ts.map +1 -0
  169. package/dist/src/wgsl/radix-hist.wgsl.js +31 -0
  170. package/dist/src/wgsl/radix-hist.wgsl.js.map +1 -0
  171. package/dist/src/wgsl/radix-scatter.wgsl.d.ts +9 -0
  172. package/dist/src/wgsl/radix-scatter.wgsl.d.ts.map +1 -0
  173. package/dist/src/wgsl/radix-scatter.wgsl.js +40 -0
  174. package/dist/src/wgsl/radix-scatter.wgsl.js.map +1 -0
  175. package/dist/src/wgsl/scan-add.wgsl.d.ts +6 -0
  176. package/dist/src/wgsl/scan-add.wgsl.d.ts.map +1 -0
  177. package/dist/src/wgsl/scan-add.wgsl.js +14 -0
  178. package/dist/src/wgsl/scan-add.wgsl.js.map +1 -0
  179. package/dist/src/wgsl/scan-block.wgsl.d.ts +8 -0
  180. package/dist/src/wgsl/scan-block.wgsl.d.ts.map +1 -0
  181. package/dist/src/wgsl/scan-block.wgsl.js +30 -0
  182. package/dist/src/wgsl/scan-block.wgsl.js.map +1 -0
  183. package/dist/src/wgsl/segmented-reduce.wgsl.d.ts +22 -8
  184. package/dist/src/wgsl/segmented-reduce.wgsl.d.ts.map +1 -1
  185. package/dist/src/wgsl/segmented-reduce.wgsl.js +84 -15
  186. package/dist/src/wgsl/segmented-reduce.wgsl.js.map +1 -1
  187. package/dist/src/wgsl/spmv-pull.wgsl.d.ts +22 -11
  188. package/dist/src/wgsl/spmv-pull.wgsl.d.ts.map +1 -1
  189. package/dist/src/wgsl/spmv-pull.wgsl.js +110 -36
  190. package/dist/src/wgsl/spmv-pull.wgsl.js.map +1 -1
  191. package/dist/tsconfig.build.tsbuildinfo +1 -1
  192. package/dist/webgpu-graph-algorithms.js +3815 -1003
  193. package/dist/webgpu-graph-algorithms.js.map +1 -1
  194. package/package.json +9 -8
  195. package/src/algorithms/components.ts +12 -16
  196. package/src/algorithms/degree.ts +58 -43
  197. package/src/algorithms/pagerank.ts +20 -18
  198. package/src/algorithms/power-iteration.ts +19 -18
  199. package/src/constants.ts +38 -8
  200. package/src/errors.ts +3 -1
  201. package/src/index.ts +14 -4
  202. package/src/kernel/dispatch.ts +18 -7
  203. package/src/kernel/kernel.ts +59 -5
  204. package/src/kernel/prelude.ts +9 -0
  205. package/src/kernel/profiler.ts +28 -4
  206. package/src/kernels.ts +356 -18
  207. package/src/layouts/calibrate.ts +187 -0
  208. package/src/layouts/force-simulation.ts +91 -23
  209. package/src/layouts/forceatlas2.ts +331 -106
  210. package/src/layouts/fruchterman-reingold.ts +255 -74
  211. package/src/layouts/model-common.ts +98 -3
  212. package/src/layouts/repulsion-grid.ts +451 -0
  213. package/src/layouts/spring-electrical.ts +257 -78
  214. package/src/memory/residency.ts +126 -20
  215. package/src/primitives/core-shape.ts +91 -4
  216. package/src/primitives/grid-pyramid.ts +221 -0
  217. package/src/primitives/grid.ts +349 -0
  218. package/src/primitives/histogram.ts +273 -0
  219. package/src/primitives/radix-sort.ts +246 -0
  220. package/src/primitives/scan.ts +197 -0
  221. package/src/primitives/segmented-reduce.ts +214 -56
  222. package/src/primitives/spmv.ts +125 -65
  223. package/src/primitives/verify.ts +249 -0
  224. package/src/types/context.ts +56 -0
  225. package/src/types/layout.ts +22 -0
  226. package/src/wgsl/counting-scatter.wgsl.ts +16 -0
  227. package/src/wgsl/fa2-attraction.wgsl.ts +98 -20
  228. package/src/wgsl/fa2-stats-finalize.wgsl.ts +22 -1
  229. package/src/wgsl/grid-cell-key.wgsl.ts +29 -0
  230. package/src/wgsl/grid-centroid-hub.wgsl.ts +28 -0
  231. package/src/wgsl/grid-centroid.wgsl.ts +28 -0
  232. package/src/wgsl/grid-downsample.wgsl.ts +27 -0
  233. package/src/wgsl/grid-far-field.wgsl.ts +97 -0
  234. package/src/wgsl/grid-near-field.wgsl.ts +128 -0
  235. package/src/wgsl/histogram.wgsl.ts +14 -0
  236. package/src/wgsl/indirect-finalize.wgsl.ts +25 -0
  237. package/src/wgsl/radix-hist.wgsl.ts +30 -0
  238. package/src/wgsl/radix-scatter.wgsl.ts +39 -0
  239. package/src/wgsl/scan-add.wgsl.ts +13 -0
  240. package/src/wgsl/scan-block.wgsl.ts +29 -0
  241. package/src/wgsl/segmented-reduce.wgsl.ts +84 -15
  242. package/src/wgsl/spmv-pull.wgsl.ts +110 -36
  243. package/dist/chunks/context-BR7fx3vR.js.map +0 -1
package/README.md CHANGED
@@ -3,20 +3,24 @@
3
3
  WebGPU-accelerated graph algorithms and layouts over the `@graphty/graph-format` snapshot, for Node
4
4
  (Dawn, through the `webgpu` npm package) and browsers (Chromium). One code base, three entry points:
5
5
 
6
- | Entry | Import | What it gives you |
7
- | ------------------------------------------ | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
8
- | `@graphty/webgpu-graph-algorithms` | the core | `createForceAtlas2`, `createAccelerator`, `GpuContext`, `degree`, `seedPositions`, `WebGpuGraphError`, `isSoftwareAdapter`, the constants (`EXACT_MAX_NODES`, `FA2_DEFAULTS`, `LAYOUT_TUNING_DEFAULTS`, ...) and the option / stats / accelerator types |
9
- | `@graphty/webgpu-graph-algorithms/node` | Node only | `createNodeGpuContext`, `probeNodeWebGpu`, `createNodeGpu` (Dawn), `dawnFlags` |
10
- | `@graphty/webgpu-graph-algorithms/browser` | browsers only | `probeBrowserWebGpu`, `requestGpuContext` |
11
-
12
- **Status: phase P3 (ForceAtlas2, exact tier).** The GPU ForceAtlas2 is usable from Node (`run()`) and from a
13
- browser frame loop (`step()` once per frame) up to `exactMaxNodes` = 32768 nodes with the default
14
- `repulsion: "auto"`, and at any size with `repulsion: "exact"` (all pairs, O(n^2) per iteration: 18.971 ms per
15
- iteration at 100k nodes / 1M edges on an RTX 4070 SUPER). The grid tier for 10^5-10^6 nodes (P4),
16
- Fruchterman-Reingold (P5) and the algorithms (P7+) follow the phase plan of `design/webgpu/webgpu-acceleration-plan.md` (monorepo root)
17
- section 13 and the interface contract `design/webgpu/plans/2026-09-14-webgpu-p0-p3-interfaces.md`; the gate
18
- record of this phase is `docs/decisions/G3.md`. There is no CPU fallback anywhere in this package: when no
19
- adapter or device exists it throws `WebGpuGraphError`.
6
+ | Entry | Import | What it gives you |
7
+ | ------------------------------------------ | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
8
+ | `@graphty/webgpu-graph-algorithms` | the core | the layouts (`createForceAtlas2`, `createFruchtermanReingold`, `createSpringElectrical`, `seedPositions`), the algorithms (`pageRank`, `personalizedPageRank`, `hits`, `eigenvectorCentrality`, `katzCentrality`, `connectedComponents`, `degree`), `createAccelerator`, `calibrateLayout`, `verifyDevice`, `GpuContext`, `WebGpuGraphError`, `isSoftwareAdapter`, the constants (`EXACT_MAX_NODES`, `FA2_DEFAULTS`, `FR_DEFAULTS`, `SE_DEFAULTS`, `LAYOUT_TUNING_DEFAULTS`, ...) and the option / stats / accelerator types |
9
+ | `@graphty/webgpu-graph-algorithms/node` | Node only | `createNodeGpuContext`, `probeNodeWebGpu`, `createNodeGpu` (Dawn), `dawnFlags` |
10
+ | `@graphty/webgpu-graph-algorithms/browser` | browsers only | `probeBrowserWebGpu`, `requestGpuContext` |
11
+
12
+ **Status: three force layouts and six algorithms, on the exact and the grid repulsion tiers.**
13
+ ForceAtlas2, Fruchterman-Reingold and ngraph's spring-electrical preset run from Node (`run()`) and from a
14
+ browser frame loop (`step()` once per frame) at any size: the default `repulsion: "auto"` runs the exact tier
15
+ up to `exactMaxNodes` = 32768 nodes (owner decision G4-D1 in `docs/decisions/G4.md`) and the grid tier above
16
+ it; `repulsion: "exact"` forces all pairs at any size (O(n^2) per iteration: 18.971 ms per iteration at 100k
17
+ nodes / 1M edges on an RTX 4070 SUPER) and `repulsion: "grid"` forces the grid tier. PageRank, personalized
18
+ PageRank, HITS, eigenvector centrality, Katz centrality and weakly connected components run as plain async
19
+ calls on the same context. The phases follow the phase plan of
20
+ `design/webgpu/webgpu-acceleration-plan.md` (monorepo root) section 13 and the interface contract
21
+ `design/webgpu/plans/2026-09-14-webgpu-p0-p3-interfaces.md`; the gate records are `docs/decisions/G<n>.md`.
22
+ There is no CPU fallback anywhere in this package: when no adapter or device exists it throws
23
+ `WebGpuGraphError`.
20
24
 
21
25
  ## Install
22
26
 
@@ -61,8 +65,8 @@ ctx.release(snapshot); // destroys the snapshot's buffers (nothing is freed by G
61
65
  ctx.dispose(); // destroys the device and lets the process exit
62
66
  ```
63
67
 
64
- Sizes above `exactMaxNodes` need `repulsion: "exact"` until the grid tier lands (P4); with the default
65
- `"auto"` the simulation rejects `load()` with `E_UNSUPPORTED { feature: "repulsion.grid" }`. The same run
68
+ Above `exactMaxNodes` the default `"auto"` picks the grid tier (spec 7.8: by n alone); `repulsion: "exact"`
69
+ keeps the all-pairs tier at any size and `repulsion: "grid"` picks the grid tier at any size. The same run
66
70
  from the command line, with a verification of the result:
67
71
 
68
72
  ```bash
@@ -85,7 +89,7 @@ if (!probe.ok) {
85
89
  throw new Error(`${probe.code}: ${probe.reason ?? ""}`); // E_NO_WEBGPU, E_NO_ADAPTER or E_SOFTWARE_ONLY
86
90
  }
87
91
  const ctx = await requestGpuContext({ adapter: probe.adapter ?? undefined });
88
- const acc = createAccelerator(ctx, { layout: { exactMaxNodes: 32768 } }); // the tuning every simulation inherits
92
+ const acc = createAccelerator(ctx, { layout: { exactMaxNodes: 4096 } }); // the tuning every simulation inherits
89
93
  const sim = acc.forceAtlas2({ seed: 1, iterationsPerStep: 1, maxInFlight: 2 });
90
94
  sim.load(snapshot, positions); // positions: your stride-3 Float32Array; NaN rows are seeded
91
95
 
@@ -146,13 +150,13 @@ names and defaults as the CPU port in `@graphty/layout`) plus the GPU tuning:
146
150
  GPU tuning (`GpuLayoutTuning`; also the `layout` field of `createAccelerator`'s options, inherited by every
147
151
  simulation the accelerator creates):
148
152
 
149
- | Option | Default | Meaning |
150
- | --------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
151
- | `repulsion` | `"auto"` | `"exact"` at any n; `"auto"` = exact iff n <= `exactMaxNodes`; `"grid"` is `E_UNSUPPORTED` until P4 |
152
- | `exactMaxNodes` | `32768` | the crossover, measured on the RTX 4070 SUPER (`docs/decisions/G3.md`); pass your own for another GPU (P4's `calibrateLayout` measures it) |
153
- | `deterministic` | `true` | fixed summation order (the exact tier is always deterministic) |
154
- | `compat` | `"paper"` | `"networkx"` reproduces NetworkX 3.4's `forceatlas2_layout` (gravity toward the origin, its accumulated swing / traction) |
155
- | `nearMax`, `gridMax2D`, `gridMax3D`, `extentFactor` | `64`, `512`, `128`, `6` | stored for the grid tier (P4) |
153
+ | Option | Default | Meaning |
154
+ | --------------------------------------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
155
+ | `repulsion` | `"auto"` | `"exact"` at any n; `"auto"` = exact iff n <= `exactMaxNodes`, else grid; `"grid"` at any n |
156
+ | `exactMaxNodes` | `32768` | the G3 value, kept by owner decision G4-D1 while the accuracy work G4-F1 leaves is open (`docs/decisions/G4.md`; the section 7.8 rule computes 1024 on the RTX 4070 SUPER); pass your own (`calibrateLayout` measures it). |
157
+ | `deterministic` | `true` | fixed summation order (the exact tier is always deterministic) |
158
+ | `compat` | `"paper"` | `"networkx"` reproduces NetworkX 3.4's `forceatlas2_layout` (gravity toward the origin, its accumulated swing / traction) |
159
+ | `nearMax`, `gridMax2D`, `gridMax3D`, `extentFactor` | `64`, `512`, `128`, `6` | stored for the grid tier (P4) |
156
160
 
157
161
  Every range error is `E_INVALID_ARGUMENT` (`gravity < 0`, `scalingRatio <= 0`, `jitterTolerance <= 0`,
158
162
  `maxIter < 1`, `settleWindow < 1`, `maxInFlight < 1`, `iterationsPerStep < 1`, `dim` not 2 or 3,
@@ -169,6 +173,150 @@ else the batch's wall time divided by its iteration count), the grid fields (`nu
169
173
  `trace`: one `{ swing, traction, speed, speedEfficiency, meanDisplacement, settledCount }` record per
170
174
  iteration of the last batch.
171
175
 
176
+ ## Fruchterman-Reingold and the spring-electrical preset
177
+
178
+ Two more force layouts ship beside ForceAtlas2, on the same exact repulsion tier and with exactly the
179
+ lifecycle the two ForceAtlas2 sections above describe: `load(snapshot, positions)`, then `run()` from Node or
180
+ `step()` once per frame, with `stats`, `settled`, `setPosition`, `setFixed`, `reheat`, `flush`, `setParams`,
181
+ `dispose` and the coalescing rule all unchanged. They differ in the force law, in their option records and in
182
+ what their trace carries. Both are also methods of an accelerator: `acc.fruchtermanReingold(options)` and
183
+ `acc.springElectrical(options)`, which pass the accelerator's `layout` tuning down the same way
184
+ `acc.forceAtlas2` does.
185
+
186
+ What they ask of the graph, and what they refuse, is ForceAtlas2's:
187
+
188
+ - an UNDIRECTED snapshot. `load()` throws `E_SNAPSHOT { reason: "directed" }` otherwise; pass
189
+ `toUndirected().snapshot`
190
+ - edge weights are ignored by both models, and both give every node the same mass rule for its whole run --
191
+ 1 for Fruchterman-Reingold, `1 + degree / 3` for the spring preset -- so neither has a `weight` or
192
+ `nodeMass` option to pass
193
+ - `positions` is the owner's stride-3 scene-unit `Float32Array` of length `3 * nodeCount`, over a plain
194
+ `ArrayBuffer`; any other length is `E_INVALID_ARGUMENT`. Rows holding a non-finite component are seeded in
195
+ place at `load()` (see `seedPositions` below), finite rows are the starting layout
196
+ - an empty snapshot (`nodeCount: 0`) loads without touching the GPU and is `settled` on arrival; `step()`
197
+ resolves immediately and the positions array is left alone
198
+ - a graph the device cannot hold is refused at `load()`, before any GPU work:
199
+ `E_TOO_LARGE { path: "positions" }` when `16 * nodeCount` bytes exceed the device's `maxBufferSize`,
200
+ `{ path: "windowed" }` when the arc arrays need more than one storage binding, and `{ path: "partials" }`
201
+ above 16,776,960 nodes
202
+ - above `exactMaxNodes` with the default `repulsion: "auto"`, `load()` throws
203
+ `E_UNSUPPORTED { feature: "repulsion.grid" }`, as ForceAtlas2 does; `repulsion: "exact"` runs all pairs at
204
+ any size
205
+
206
+ The GPU tuning table above applies to both, with one exception: `compat` is read by ForceAtlas2 alone and
207
+ changes nothing here.
208
+
209
+ ### Fruchterman-Reingold
210
+
211
+ The spring model of the original paper: every pair of nodes repels with `k^2 / d`, every edge pulls with
212
+ `d^2 / k`, where `k` is the ideal edge length, and a falling temperature caps how far a node may move in a
213
+ single iteration.
214
+
215
+ ```ts
216
+ import { createFruchtermanReingold } from "@graphty/webgpu-graph-algorithms";
217
+
218
+ // positions: the owner's stride-3 scene-unit array; NaN rows are seeded at load(), finite rows are kept
219
+ const positions = new Float32Array(3 * snapshot.nodeCount).fill(Number.NaN);
220
+
221
+ const sim = createFruchtermanReingold(ctx, { iterations: 50, cooling: "linear", seed: 42, dim: 2 });
222
+ sim.load(snapshot, positions); // an undirected snapshot
223
+ const stats = await sim.run({ batch: 8 }); // stops at `iterations` or when the layout settles
224
+ console.log(sim.iterationsDone, sim.settled, stats.temperature, stats.meanDisplacement);
225
+
226
+ sim.dispose(); // then ctx.release(snapshot) and ctx.dispose() when the graph goes away
227
+ ```
228
+
229
+ | Option | Default | Meaning |
230
+ | --------------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
231
+ | `k` | `null` | the ideal edge length. `null` (and `0`, and `NaN`) means `1 / sqrt(nodeCount)`, resolved at `load()`; anything else must be a finite number > 0 |
232
+ | `iterations` | `50` | an integer >= 0 and the budget `run()` stops at (`0` is settled on arrival). Under `cooling: "adaptive"` it is only a cap, and an option record that never set it gets `10000` instead |
233
+ | `cooling` | `"linear"` | `"linear"`: the temperature falls from `0.1` to `0` across `iterations`, so a run lasts the whole budget. `"adaptive"`: Yifan Hu's step control -- the temperature shrinks by `0.9` whenever the total force energy rose and grows by `1 / 0.9` after five consecutive falls -- so a run ends when the layout stops moving rather than at a fixed count |
234
+ | `fixed` | `null` | which nodes are pinned, applied at `load()`: a graph-format `NodeMask`, the name of a bool node column, or `null`, which takes the column with role `fixed` when the snapshot has one and pins nothing otherwise. Not a live option: `setParams({ fixed })` is `E_INVALID_ARGUMENT`, use `setFixed(mask)` |
235
+ | `dim`, `scale`, `center`, `seed` | `2`, `1`, `[0, 0, 0]`, `null` | as ForceAtlas2 |
236
+ | `settleThreshold`, `settleWindow`, `iterationsPerStep`, `maxInFlight` | `0.001`, `10`, `1`, `2` | as ForceAtlas2 |
237
+
238
+ `FR_DEFAULTS` is the frozen record of those defaults (`center` and `seed` are the shared ones and are not in
239
+ it). Every range error is `E_INVALID_ARGUMENT`, and `maxInFlight` cannot change after creation.
240
+
241
+ `sim.stats` is `FruchtermanReingoldStats`: everything in `LayoutStatsBase` -- `iteration`,
242
+ `meanDisplacement`, `rmsRadius`, `layoutRadius`, `centroid`, `repulsionTier` (`"exact"`), `msPerIteration`
243
+ and the grid fields, which are `null` -- plus `temperature`, the schedule's value for the last completed
244
+ iteration, and `trace`: one `{ temperature, meanDisplacement, settledCount }` record per iteration of the
245
+ last batch.
246
+
247
+ ### The spring-electrical preset
248
+
249
+ ngraph.forcelayout's model with ngraph's own constants: Hooke springs along the edges, Coulomb repulsion
250
+ between every pair, and a velocity integrated with drag. A node's mass is `1 + degree / 3`, as ngraph
251
+ assigns it.
252
+
253
+ ```ts
254
+ import { createSpringElectrical } from "@graphty/webgpu-graph-algorithms";
255
+
256
+ const sim = createSpringElectrical(ctx, { springLength: 10, timeStep: 0.5, seed: 42 });
257
+ sim.load(snapshot, positions);
258
+ const stats = await sim.run({ maxIter: 1000, batch: 8 }); // pass maxIter: this model has no iteration count
259
+ console.log(sim.iterationsDone, sim.settled, stats.kineticEnergy);
260
+
261
+ sim.dispose();
262
+ ```
263
+
264
+ | Option | Default | Meaning |
265
+ | --------------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
266
+ | `springLength` | `10` | the rest length of an edge; a finite number > 0 |
267
+ | `springCoefficient` | `0.8`, size-scaled | Hooke's constant. Left out or `null` it is ngraph's `0.8` times `min(1, 300 / nodeCount)`, applied once `nodeCount` is known at `load()`; a number you pass is used as given at every size |
268
+ | `gravity` | `-12`, size-scaled | ngraph's Coulomb constant, where NEGATIVE repels -- this is not ForceAtlas2's centre gravity. Left out or `null` it is ngraph's `-12` times the same factor. Any finite number is accepted |
269
+ | `dragCoefficient` | `0.9` | velocity damping; a finite number >= 0 |
270
+ | `timeStep` | `0.5` | the integrator's step; a finite number > 0 |
271
+ | `dim`, `scale`, `center`, `seed` | `2`, `1`, `[0, 0, 0]`, `null` | as ForceAtlas2 |
272
+ | `settleThreshold`, `settleWindow`, `iterationsPerStep`, `maxInFlight` | `0.001`, `10`, `1`, `2` | as ForceAtlas2 |
273
+
274
+ The size scaling exists because ngraph's constants were tuned for graphs of a few hundred nodes: on tens of
275
+ thousands the per-node forces are large enough that every node moves at the unit speed clamp and the layout
276
+ never comes to rest. `SE_DEFAULTS` is the frozen record of the unscaled constants.
277
+
278
+ This model has no iteration-count option, so `run()` without `maxIter` has no budget at all and returns only
279
+ when the layout settles. Pass `run({ maxIter })` unless that is what you want.
280
+
281
+ `sim.stats` is `SpringElectricalStats`: `LayoutStatsBase` plus `kineticEnergy`, and a `trace` of
282
+ `{ kineticEnergy, meanDisplacement, settledCount }`. The energy lags the positions by one iteration -- the
283
+ iteration that follows an integrate is the one that folds its energy -- so the first record after `load()`
284
+ carries `0` and record i carries the energy of iteration i - 1.
285
+
286
+ Both models appear in the Performance section below as row T-14 and in the `layout-fr` benchmark group.
287
+
288
+ ## The `seedPositions` helper
289
+
290
+ Every simulation seeds the unplaced rows of your positions array at `load()`, so you never have to call this.
291
+ It is public for the case where you want the same starting layout without a simulation -- to draw the graph
292
+ before the first frame lands, to reproduce a run on the CPU, or to place a subgraph the same way twice.
293
+
294
+ ```ts
295
+ import { fromEdgeArrays } from "@graphty/graph-format";
296
+ import { seedPositions } from "@graphty/webgpu-graph-algorithms";
297
+
298
+ const s = fromEdgeArrays({ directed: false, nodeCount: 3, src: new Uint32Array([0]), dst: new Uint32Array([1]) });
299
+ const positions = new Float32Array(3 * s.nodeCount).fill(Number.NaN);
300
+
301
+ seedPositions(s, positions, 42, 2, 1, null, "fa2");
302
+ // every row now holds x and y in [-1, 1) and z = 0; seed 42 gives this same array every time
303
+ ```
304
+
305
+ The arguments are positional and all required: the snapshot, the owner's stride-3 scene-unit array (modified
306
+ in place, length `3 * nodeCount`), the LCG seed (`0` or `null` draws a random one, the CPU port's quirk kept
307
+ bit for bit), `dim` (`2` or `3`), the scene `scale` (> 0), the scene `center` (an `ArrayLike<number>`, or
308
+ `null` for the origin), and the draw range: `"fa2"` is `[-1, 1)` in layout units, `"fr"` is `[0, 1)`.
309
+ ForceAtlas2 and the spring preset seed with `"fa2"`, Fruchterman-Reingold with `"fr"`.
310
+
311
+ A row counts as unseeded when ANY of its first `dim` components is not finite, and only its non-finite
312
+ components are written -- a finite component is never changed and a fully finite row is never touched. In 2D
313
+ the third component of a row being seeded is set to `center[2]`. The draw box depends on what is already
314
+ there: when no row is fully finite the draw is the plain range per axis, and otherwise it is the `[min, max]`
315
+ box of the finite components, per axis, so new nodes land among the ones already placed instead of around the
316
+ origin. No random number is drawn at all when nothing needs seeding. A bad `dim`, a `positions` length other
317
+ than `3 * nodeCount`, a non-positive `scale`
318
+ or a non-finite `center` component is `E_INVALID_ARGUMENT`; the function returns nothing.
319
+
172
320
  ## Acquisition
173
321
 
174
322
  ### Node
@@ -243,10 +391,218 @@ ctx.dispose(); // destroys the device and lets the process exit
243
391
  In a browser, `import { requestGpuContext } from "@graphty/webgpu-graph-algorithms/browser"` and `const ctx = await
244
392
  requestGpuContext();` replace the first import and line; the rest is identical.
245
393
 
394
+ ## Centrality and components
395
+
396
+ Six algorithm functions run on the device. Every one has the same shape -- `await fn(ctx, snapshot, options?)`
397
+ -- and every array that comes back is indexed by the snapshot's NODE INDEX (`0 .. nodeCount - 1`), never by
398
+ node id; the snapshot's id map turns an index back into the id it was built from. Every score result carries
399
+ `precision: "f32"`, so a reader can label what it is looking at.
400
+
401
+ The context uploads what an algorithm needs the first time it sees a snapshot and keeps it until
402
+ `ctx.release(snapshot)`: the CSR core for PageRank, HITS, eigenvector centrality and connected components,
403
+ the reverse adjacency for PageRank, HITS and Katz, the edge list for connected components. A second call on
404
+ the same snapshot pays no upload.
405
+
406
+ Each one is also a method of the object `createAccelerator(ctx)` returns -- `acc.pageRank(s, options)` and so
407
+ on, with `connectedComponents` answering to `weaklyConnectedComponents` as well -- which is how
408
+ `@graphty/algorithms` reaches them when a GPU accelerator is injected into it.
409
+
410
+ The PageRank example below is complete; the examples after it reuse its context `ctx` and its snapshot `s`
411
+ rather than repeating the setup.
412
+
413
+ Options every one of them accepts on top of its own (`GpuRunOptions`):
414
+
415
+ | Option | Default | Meaning |
416
+ | ------------ | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
417
+ | `dest` | unset | a preallocated result of exactly `nodeCount` elements over a plain `ArrayBuffer`: a `Float32Array` for scores, a `Uint32Array` for labels. Any other type or length is `E_INVALID_ARGUMENT { argument: "dest" }`; `hits` fills it with the hubs |
418
+ | `signal` | unset | an `AbortSignal`, checked before the first submit and after every batch; `E_ABORTED { batchId }` |
419
+ | `onProgress` | unset | `(done, total)` after every batch. `total` is `maxIterations`, `2 * maxIterations` for `hits` (it runs two chains) and `1` for `connectedComponents`, which reports once when it finishes |
420
+
421
+ A graph too large for the device is refused before any GPU work: when the arc arrays need more than one
422
+ storage binding they would need a windowed upload, which only `degree` executes, and each of these six throws
423
+ `E_TOO_LARGE { needed, limit, path: "windowed" }` instead. A node count whose dispatch does not fit
424
+ the device's 2D workgroup grid is `E_TOO_LARGE { path: "dispatch" }`, and scratch that does not fit the
425
+ device's budget is `E_OUT_OF_MEMORY { requested, resident, label }`.
426
+
427
+ ### PageRank
428
+
429
+ PageRank is the stationary distribution of a surfer who follows an out-arc with probability `dampingFactor`
430
+ and teleports to a uniformly random node otherwise. `pageRank` runs NetworkX's iteration on the device: rank
431
+ is pulled along the reverse adjacency, out-weights normalise the push, and the mass sitting on nodes with no
432
+ out-arcs is redistributed through the teleport vector every iteration.
433
+
434
+ ```ts
435
+ import { fromEdgeArrays } from "@graphty/graph-format";
436
+ import { pageRank } from "@graphty/webgpu-graph-algorithms";
437
+ import { createNodeGpuContext } from "@graphty/webgpu-graph-algorithms/node";
438
+
439
+ const ctx = await createNodeGpuContext();
440
+ const s = fromEdgeArrays({
441
+ directed: true,
442
+ nodeCount: 4,
443
+ src: new Uint32Array([0, 1, 2, 3]), // a tail into a 3-cycle: 0 -> 1 -> 2 -> 3 -> 1
444
+ dst: new Uint32Array([1, 2, 3, 1]),
445
+ });
446
+
447
+ const r = await pageRank(ctx, s, { dampingFactor: 0.85, tolerance: 1e-6 });
448
+ console.log(r.scores); // Float32Array(4), index-aligned, summing to 1; node 0 scores lowest
449
+ console.log(r.iterations, r.converged, r.danglingMass);
450
+
451
+ ctx.release(s); // when the graph goes away; ctx.dispose() at the end of the program
452
+ ```
453
+
454
+ | Option | Default | Meaning |
455
+ | --------------- | ------- | -------------------------------------------------------------------------------------------------------------------- |
456
+ | `dampingFactor` | `0.85` | the probability of following an arc; the rest teleports |
457
+ | `maxIterations` | `100` | an integer >= 1; anything else is `E_INVALID_ARGUMENT` |
458
+ | `tolerance` | `1e-6` | converged when the L1 change of the whole vector falls below `tolerance * nodeCount` (NetworkX's rule) |
459
+ | `weighted` | `true` | arc weights are the transition mass; `false` weighs every arc 1. A snapshot with no weights is unweighted either way |
460
+
461
+ `GpuPageRankResult` comes back: `scores` (one f32 per node, summing to 1 up to f32 rounding), `iterations`
462
+ (the FIRST iteration whose delta fell below the threshold, not the batch boundary the run stopped at, and
463
+ `maxIterations` when it never did), `converged`, `danglingMass` (the rank mass the last iteration found on
464
+ nodes with no positive out-weight and redistributed) and `precision`.
465
+
466
+ Directed or undirected both work: an undirected snapshot carries both directions, so the pull and the
467
+ normaliser see the same arcs, and the scores differ from the directed form's as they should. `nodeCount: 0`
468
+ returns an empty `scores` with `iterations: 0`, `converged: true` and does no GPU work. A graph with no arcs
469
+ at all makes every node dangling, so the teleport vector is already the fixed point: the call returns
470
+ `1 / nodeCount` everywhere with `iterations: 0` and `danglingMass: 1`. An isolated node inside a larger graph
471
+ is one dangling node -- it keeps its teleport share and its share of the redistributed mass -- and so is a
472
+ node whose out-arcs all weigh zero under `weighted: true`.
473
+
474
+ ### Personalized PageRank
475
+
476
+ The same iteration with the uniform teleport vector replaced by yours, so the walk restarts where you say:
477
+
478
+ ```ts
479
+ import { personalizedPageRank } from "@graphty/webgpu-graph-algorithms";
480
+
481
+ const bias = new Float32Array(s.nodeCount);
482
+ bias[0] = 1; // restart at node 0 only
483
+ const r = await personalizedPageRank(ctx, s, bias, { dampingFactor: 0.85 });
484
+ console.log(r.scores); // mass concentrated on what node 0 reaches
485
+ ```
486
+
487
+ `personalization` is a `Float32Array` of one finite non-negative number per node, not all zero, and it is
488
+ normalised to sum 1 on the host before the run -- so unnormalised weights are fine. A wrong length, a
489
+ negative or non-finite entry, or a zero total is `E_INVALID_ARGUMENT { argument: "personalization" }`.
490
+ Options, result and edge cases are PageRank's, except that a graph with no arcs returns the normalised
491
+ personalization vector rather than `1 / nodeCount`.
492
+
493
+ ### HITS
494
+
495
+ HITS scores each node twice: as a hub (it points at good authorities) and as an authority (good hubs point at
496
+ it). `hits` runs the CPU package's recurrence -- `a(i) = A^T norm(h(i-1))` and `h(i) = A norm(a(i-1))` from
497
+ uniform seeds -- as two interleaved chains on the device, then sum-normalises both vectors once on the host.
498
+
499
+ ```ts
500
+ import { hits } from "@graphty/webgpu-graph-algorithms";
501
+
502
+ const cycle = fromEdgeArrays({
503
+ directed: true,
504
+ nodeCount: 2,
505
+ src: new Uint32Array([0, 1]),
506
+ dst: new Uint32Array([1, 0]),
507
+ });
508
+ const r = await hits(ctx, cycle);
509
+ console.log(r.hubs, r.authorities); // [0.5, 0.5] and [0.5, 0.5]
510
+ ctx.release(cycle);
511
+ ```
512
+
513
+ The options are `maxIterations` (`100`), `tolerance` (`1e-6`) and `weighted` (`true`), with PageRank's
514
+ meanings. `GpuHitsResult` carries `hubs` and `authorities` (both f32, index-aligned, each summing to 1),
515
+ `iterations` (the larger of the two chains'), `converged` (both chains) and `precision`. `dest` receives the
516
+ hubs; the authorities always come back in a fresh array.
517
+
518
+ Direction is the point of the algorithm: an undirected snapshot has a symmetric adjacency, so the two vectors
519
+ come out the same. `nodeCount: 0` returns two empty arrays with `iterations: 0` and `converged: true`; a graph
520
+ with no arcs returns all zeros in both, since there is nothing to be a hub of.
521
+
522
+ ### Eigenvector centrality
523
+
524
+ A node is central when central nodes point at it. The scores are the principal eigenvector of the adjacency
525
+ matrix, found by power iteration over the forward adjacency and L2-normalised once at the end.
526
+
527
+ ```ts
528
+ import { eigenvectorCentrality } from "@graphty/webgpu-graph-algorithms";
529
+
530
+ const r = await eigenvectorCentrality(ctx, s, { maxIterations: 100, tolerance: 1e-6 });
531
+ console.log(r.scores, r.converged); // f32, index-aligned, L2 norm 1
532
+ ```
533
+
534
+ The options are `maxIterations` (`100`), `tolerance` (`1e-6`) and `weighted` (`true`); convergence is
535
+ PageRank's L1 rule, a delta below `tolerance * nodeCount`. It returns `GpuScoresResult` -- `scores`,
536
+ `iterations`, `converged`, `precision` -- which is PageRank's result without `danglingMass`.
537
+
538
+ Directed or undirected both work. When a graph's components have different spectral radii the power iteration
539
+ converges on the dominant component's eigenvector and the others fall toward zero; that is a property of the
540
+ method, not of this implementation. `nodeCount: 0` is empty; a graph with no arcs, and an isolated node in a
541
+ larger graph, scores 0.
542
+
543
+ ### Katz centrality
544
+
545
+ Katz counts every walk that ends at a node, discounted by its length: `x = alpha * A^T x + beta`, so a walk of
546
+ length L contributes `alpha^L`. Unlike eigenvector centrality it hands every node the floor `beta`, which
547
+ keeps a node with no incoming arcs from scoring zero.
548
+
549
+ ```ts
550
+ import { katzCentrality } from "@graphty/webgpu-graph-algorithms";
551
+
552
+ const r = await katzCentrality(ctx, s, { alpha: 0.1, beta: 1 });
553
+ console.log(r.scores); // f32, index-aligned, L2-normalised
554
+ ```
555
+
556
+ | Option | Default | Meaning |
557
+ | --------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
558
+ | `alpha` | `0.1` | the attenuation per step. Convergence needs it below the reciprocal of the largest eigenvalue; the package checks only that it is finite |
559
+ | `beta` | `1` | the constant every node is given each iteration |
560
+ | `maxIterations` | `100` | as above |
561
+ | `tolerance` | `1e-6` | as above |
562
+ | `weighted` | `true` | as above |
563
+
564
+ A non-finite `alpha` or `beta` is `E_INVALID_ARGUMENT`. The result is `GpuScoresResult`, L2-normalised on the
565
+ host with no per-iteration normaliser on the device. Walks arrive along the reverse adjacency, so direction
566
+ matters. `nodeCount: 0` is empty; with no arcs every node holds `beta` alone, which after normalisation is
567
+ `1 / sqrt(nodeCount)` each.
568
+
569
+ ### Connected components
570
+
571
+ `connectedComponents` labels the WEAK components: it walks the edge list, so an edge joins its two endpoints
572
+ whether or not the snapshot is directed, and the directed and undirected forms of one edge set give identical
573
+ labels. The kernel is GAP's Afforest -- sampled link rounds, then each edge once, then a final compress.
574
+
575
+ ```ts
576
+ import { connectedComponents } from "@graphty/webgpu-graph-algorithms";
577
+
578
+ const g = fromEdgeArrays({
579
+ directed: false,
580
+ nodeCount: 4,
581
+ src: new Uint32Array([0, 1]), // the path 0 - 1 - 2, with node 3 on its own
582
+ dst: new Uint32Array([1, 2]),
583
+ });
584
+ const r = await connectedComponents(ctx, g);
585
+ console.log(r.labels); // Uint32Array [0, 0, 0, 1]
586
+ console.log(r.count); // 2
587
+ console.log(r.groups()); // [Uint32Array [0, 1, 2], Uint32Array [3]]
588
+ ctx.release(g);
589
+ ```
590
+
591
+ `labels` holds one label per node index. By default they are renumbered dense `0 .. count - 1` in first-seen
592
+ index order, which makes them identical to the labels `@graphty/algorithms` produces for the same graph;
593
+ `renumber: false` returns the raw root indices instead -- the same partition with arbitrary label values, and
594
+ the same `count`. `groups()` builds the member lists on the first call and caches them, index-aligned with
595
+ the labels in that same first-seen order.
596
+
597
+ `renumber` (default `true`) is the only option besides the shared three. `nodeCount: 0` gives empty labels,
598
+ `count: 0` and no GPU work; a graph with no arcs gives every node its own block, so `labels[v] === v` and
599
+ `count === nodeCount`, which is what an isolated node gets inside a larger graph too. Self-loops and repeated
600
+ edges change nothing: the link step is idempotent.
601
+
246
602
  ## Errors
247
603
 
248
604
  Every condition the package detects itself is a `WebGpuGraphError` with a stable `code`
249
- (`E_NO_WEBGPU`, `E_NO_ADAPTER`, `E_NO_DEVICE`, `E_SOFTWARE_ONLY`, `E_DEVICE_LOST`, `E_DISPOSED`,
605
+ (`E_NO_WEBGPU`, `E_NO_ADAPTER`, `E_NO_DEVICE`, `E_SOFTWARE_ONLY`, `E_DEVICE_LOST`, `E_DEVICE_INCORRECT`, `E_DISPOSED`,
250
606
  `E_VALIDATION`, `E_SHADER_COMPILE`, `E_OUT_OF_MEMORY`, `E_TOO_LARGE`, `E_UNSUPPORTED`,
251
607
  `E_INVALID_ARGUMENT`, `E_SNAPSHOT`, `E_RELEASED`, `E_NOT_LOADED`, `E_ABORTED`) and frozen `details`.
252
608
  `isWebGpuGraphError(x)` and `hasErrorCode(x, code)` are structural brand checks, so they survive two copies
@@ -255,6 +611,38 @@ of the package. The graph-format codes `E_GPU_INELIGIBLE`, `E_UNKNOWN_NODE`, `E_
255
611
  released rejects its next `step()` with `E_RELEASED`; a lost device disposes every simulation and rejects
256
612
  every pending `step()` with `E_DEVICE_LOST` (create a new context from a fresh adapter and `load()` again).
257
613
 
614
+ ## The device self-check
615
+
616
+ Some GPU drivers return wrong answers rather than failing. On Windows over the Microsoft Basic Render Driver,
617
+ compute shaders that synchronise across a workgroup produce silently incorrect results, which would make every
618
+ number this package computes there unreliable with no error anywhere.
619
+
620
+ So the first algorithm or layout you run on a context asks the device for an answer this package already knows --
621
+ an exclusive scan across 33 workgroups of known numbers (8,193 words on a 256-lane device), verified word by word
622
+ on the host -- and refuses a device that gets it wrong with `E_DEVICE_INCORRECT`, naming the first wrong word,
623
+ what belonged there and the adapter's description string. It runs once per device and is remembered: 14-20 ms
624
+ the first time, nearly all of it compiling the two scan pipelines that any scan-using algorithm would compile
625
+ anyway, and 0.6-1.0 ms of work under that (measured on Dawn over lavapipe and over an RTX 4070 SUPER). It cannot be switched off: a flag for it would be off in
626
+ somebody's production build, which is the silent wrong answers walking back in.
627
+
628
+ Choosing the processor instead is your decision, not the package's, so nothing falls back. To ask before you
629
+ commit work to a device, call it yourself:
630
+
631
+ ```ts
632
+ import { verifyDevice } from "@graphty/webgpu-graph-algorithms";
633
+
634
+ const check = await verifyDevice(ctx); // memoised: the algorithms below reuse this result
635
+ if (!check.ok) {
636
+ // check.mismatch names the first wrong word; check.description is the adapter string that identifies the driver
637
+ runOnTheCpuInstead();
638
+ }
639
+ ```
640
+
641
+ It checks one property -- that values crossing a workgroup barrier, and block totals crossing dispatches of one
642
+ compute pass, survive. A device that gets that right and gets atomics or float rounding wrong still passes. It is
643
+ a refusal mechanism, not a certificate of correctness; `docs/decisions/device-self-check.md` records what it
644
+ covers, what it does not, and why the package refuses such a device rather than computing around it.
645
+
258
646
  ## Benchmarks
259
647
 
260
648
  ```bash
@@ -277,10 +665,15 @@ the profiler reports; every rung starts with an untimed clock warm-up burst, bec
277
665
  SM clock at its idle 210 MHz under sparse sub-millisecond dispatches and the kernels then measure 4-15x slower),
278
666
  `pagerank` (T-8), `wcc` (T-9), `layout-fr` (T-14: `step(1)` of `createFruchtermanReingold` and `createSpringElectrical`
279
667
  at 10k and at 100k with `repulsion: "exact"`, the same two rows per model and rung as `layout-exact`, tagged `fr` /
280
- `se`). The Chromium number of T-5 comes from the `bench`-tagged browser test (`GRAPHTY_BROWSER_GPU=nvidia node
281
- scripts/run-browser-project.js`), which appends its session through the Vitest commands bridge. `exactMaxNodes` is
282
- re-fixed from the ladder by the rule of plan section 7.8 (the largest rung under 4 ms per iteration, rounded down to a
283
- power of two; `benchmarks/layout-exact.bench.ts` `exactMaxNodesFromLadder`).
668
+ `se`), `layout-grid` (T-6 and T-7: `step(1)` of the grid tier on the grid ladder 32k / 65k / 100k / 262k / 1M in 2D
669
+ and in 3D, the same two rows per rung tagged `grid` with the dimension, plus the `fa2-attraction` pass of the 1M 2D
670
+ iteration from the profiler, and the exact ladder's 1k / 4k / 8k / 16k rungs in 2D for the crossover re-check). The Chromium numbers of T-5 (10k on the exact tier, 100k on the grid tier) come from the
671
+ `bench`-tagged browser test (`GRAPHTY_BROWSER_GPU=nvidia node scripts/run-browser-project.js`), which appends its
672
+ session through the Vitest commands bridge. `exactMaxNodes` is re-fixed from the ladder by the rule of plan section 7.8
673
+ (the largest rung under 4 ms per iteration and not slower than the grid tier at the same n -- the `layout-grid` 2D rows,
674
+ which the group also records at the exact ladder's 1k / 4k / 8k / 16k rungs so the clause has a row at every rung --
675
+ rounded down to a power of two; `benchmarks/layout-exact.bench.ts` `exactMaxNodesFromLadder`); `calibrateLayout(ctx)`
676
+ measures the same crossover on a consumer's own device.
284
677
 
285
678
  ## Performance
286
679
 
@@ -291,18 +684,21 @@ are the T-table of plan section 10.4. A missed target is re-fixed by a recorded
291
684
 
292
685
  ### The dev box (nvidia-lovelace-driver580)
293
686
 
294
- Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-20T01:06:48.656Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z.
295
-
296
- | Id | What | Target | Measured |
297
- | ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ---------------------- |
298
- | T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 5.980 ms; 125.738 ms |
299
- | T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 0.870 ms |
300
- | T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 0.181 ms |
301
- | T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.586 ms; 1.053 ms |
302
- | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.400 ms (0.726 ms) |
303
- | T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 17.204 ms; 198.645 ms |
304
- | T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 145.970 ms (12.613 ms) |
305
- | T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.617 ms; 16.367 ms |
687
+ Measured on nvidia-lovelace-driver580 (NVIDIA: 580.173.02 580.173.2.0), session 2026-09-20T01:06:48.656Z, medians of 5 runs; Chromium: nvidia / lovelace (nvidia-lovelace-driver0, the description is redacted by Chromium), session 2026-09-16T02:18:11.896Z. The T-5 (100k), T-6 and T-7 rows are from the later session 2026-09-21T06:15:17.027Z (the file's last session, the current baseline, the one the crossover re-check reads; its Chromium session 2026-09-21T05:48:14.190Z), whose other rows are within 1.25x of this table's.
688
+
689
+ | Id | What | Target | Measured |
690
+ | ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- | ---------------------------- |
691
+ | T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 5.980 ms; 125.738 ms |
692
+ | T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 0.870 ms |
693
+ | T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 0.181 ms |
694
+ | T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.586 ms; 1.053 ms |
695
+ | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.400 ms (0.726 ms) |
696
+ | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 100k on the grid tier, Chromium (Node in brackets) | <= 12 ms | 3.700 ms (2.087 ms) |
697
+ | T-6 | ForceAtlas2 grid tier, GPU time per iteration (profiler) at 100k 2D; at 1M 2D; at 100k 3D | <= 10 ms; <= 100 ms; <= 20 ms | 0.634 ms; 5.414 ms; 1.306 ms |
698
+ | T-7 | Attraction gather (the `fa2-attraction` pass of the grid tier), GPU time per iteration (profiler) at 1M / 10M | <= 15 ms | 1.493 ms |
699
+ | T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 17.204 ms; 198.645 ms |
700
+ | T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 145.970 ms (12.613 ms) |
701
+ | T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.617 ms; 16.367 ms |
306
702
 
307
703
  Three rows miss their target in this session: the 1M / 10M upload (125.7 ms against 100 ms, the open owner decision of
308
704
  `docs/decisions/G1.md` section 7), the empty-submit round trip (0.181 ms against 0.1 ms: the row is measured after the
@@ -336,23 +732,28 @@ GPU time per iteration in the last batch).
336
732
 
337
733
  The first run of the GPU lane (`gpu.yml`, graphty-monorepo run 35316416067, 2026-09-18) on a machine.dev T4 -- one Tesla
338
734
  T4 (16 GB), 4 vCPU of a Xeon Platinum 8259CL, driver 580.126.20 -- wrote this baseline; `scripts/bench-compare.js` fails
339
- a later run of the lane whose median exceeds 3x these figures. The T-table targets were set on the dev box; the T4 meets
340
- T-4 and T-5 and misses T-1 (both uploads), T-2 and T-3, which is the class difference of a datacentre card behind a
341
- cloud vCPU (host-side copies and submit latency), not a regression: the exact tier's `ms / iteration` is 1.7x the
342
- RTX 4070 SUPER's at 10k and 3.0x at 65k.
343
-
344
- Measured on gpu-linux-t4 (NVIDIA: 580.126.20 580.126.20.0), session 2026-09-20T02:29:33.210Z (run 35483512705), medians of 5 runs; Chromium: nvidia / turing (nvidia-turing-driver0, the description is redacted by Chromium), session 2026-09-20T02:28:10.676Z.
345
-
346
- | Id | What | Target | Measured |
347
- | ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
348
- | T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 14.338 ms; 256.418 ms |
349
- | T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 2.957 ms |
350
- | T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 1.296 ms |
351
- | T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.972 ms; 1.953 ms |
352
- | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.600 ms (1.359 ms) |
353
- | T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 45.461 ms; 1092.799 ms |
354
- | T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 293.079 ms (28.544 ms) |
355
- | T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.942 ms; 54.232 ms (the spring preset 1.051 ms; 60.706 ms) |
735
+ a later run of the lane whose median AND minimum both exceed 1.35x the best figures this file has ever held, by at
736
+ least 2.5 ms. The T-table targets were set on the dev box; the T4 meets
737
+ T-4, T-5 and T-6 and misses T-1 (both uploads), T-2, T-3 and T-7 (the 1M attraction pass of the grid tier, 18.165 ms
738
+ against 15 ms), which is the class difference of a datacentre card behind a cloud vCPU (host-side copies and submit
739
+ latency), not a regression: the exact tier's `ms / iteration` is 1.7x the RTX 4070 SUPER's at 10k and 3.0x at 65k, and
740
+ the grid tier's is 2.8x at 100k 2D and 5.7x at 1M 2D while the attraction pass alone is 12.2x.
741
+
742
+ Measured on gpu-linux-t4 (NVIDIA: 580.126.20 580.126.20.0), session 2026-09-20T02:29:33.210Z (run 35483512705), medians of 5 runs; Chromium: nvidia / turing (nvidia-turing-driver0, the description is redacted by Chromium), session 2026-09-20T02:28:10.676Z. The T-5 (100k), T-6 and T-7 rows are from the later session 2026-09-22T20:21:53.814Z (GPU lane run 35775999450 on commit 1d2d4dcf, the file's last session and the current baseline of this class; its Chromium session 2026-09-22T20:19:30.117Z).
743
+
744
+ | Id | What | Target | Measured |
745
+ | ---- | ---------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------- |
746
+ | T-1 | Upload of the 100k / 1M weighted hot prefix (16.4 MB); 1M / 10M (164 MB) | <= 10 ms; <= 100 ms | 14.338 ms; 256.418 ms |
747
+ | T-2 | `degree` + 400 KB readback at 100k (core resident), Node | <= 2 ms | 2.957 ms |
748
+ | T-3 | Empty submit + 4-byte `readU32` round trip, Dawn | <= 0.1 ms | 1.296 ms |
749
+ | T-4 | ForceAtlas2 exact tier, GPU time per iteration (profiler) at 10k; at 16k | <= 1 ms; <= 2 ms | 0.972 ms; 1.953 ms |
750
+ | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 10k, Chromium (Node in brackets) | <= 6 ms | 2.600 ms (1.359 ms) |
751
+ | T-5 | ForceAtlas2 per-frame cost, `step(1)` + the 12n readback at 100k on the grid tier, Chromium (Node in brackets) | <= 12 ms | 8.100 ms (5.314 ms) |
752
+ | T-6 | ForceAtlas2 grid tier, GPU time per iteration (profiler) at 100k 2D; at 1M 2D; at 100k 3D | <= 10 ms; <= 100 ms; <= 20 ms | 1.790 ms; 31.057 ms; 3.954 ms |
753
+ | T-7 | Attraction gather (the `fa2-attraction` pass of the grid tier), GPU time per iteration (profiler) at 1M / 10M | <= 15 ms | 18.165 ms -- MISSED: 21 % over the target on this card (1.493 ms on the dev box; G4-F16 of docs/decisions/G4.md) |
754
+ | T-8 | PageRank, 100 iterations, wall end to end including the upload, at 100k / 1M; at 1M / 10M | <= 150 ms; <= 1.5 s | 45.461 ms; 1092.799 ms |
755
+ | T-9 | Weakly connected components (Afforest), wall end to end including the upload and the label readback, at 1M / 10M (100k / 1M in brackets) | <= 100 ms | 293.079 ms (28.544 ms) |
756
+ | T-14 | Fruchterman-Reingold exact tier, GPU time per iteration (profiler) at 10k; at 100k (`repulsion: "exact"`) | recorded | 0.942 ms; 54.232 ms (the spring preset 1.051 ms; 60.706 ms) |
356
757
 
357
758
  The exact curve (the `layout-exact` group: 2D, E = 10n, seeded G(n, m), one simulation per rung; ms / iteration from the profiler):
358
759