faster-diffbloch 0.1.4__tar.gz → 0.1.6__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/PKG-INFO +37 -36
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/README.md +36 -35
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/pyproject.toml +1 -1
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/__init__.py +1 -1
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/.gitignore +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/LICENSE +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/backend.py +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/builder.py +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/cli.py +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/batch_cgemm.c +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/batch_cgemm.h +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/batch_cgemm.metal +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/bridge_lib.c +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/metal_batch_cgemm.m +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/native_scattering.c +0 -0
- {faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/native_scattering.h +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.5
|
|
2
2
|
Name: faster-diffbloch
|
|
3
|
-
Version: 0.1.
|
|
3
|
+
Version: 0.1.6
|
|
4
4
|
Summary: Drop-in Metal GPU (macOS) and optimized CPU (macOS/Linux) acceleration for diffBloch
|
|
5
5
|
Project-URL: Homepage, https://godofecht.github.io/diffFlow/
|
|
6
6
|
Project-URL: Documentation, https://godofecht.github.io/diffFlow/
|
|
@@ -35,7 +35,7 @@ Description-Content-Type: text/markdown
|
|
|
35
35
|
|
|
36
36
|
Drop-in Apple Silicon Metal GPU and optimized CPU acceleration for [diffBloch](https://diffbloch.com) electron crystallography structure refinement.
|
|
37
37
|
|
|
38
|
-
|
|
38
|
+
**2.59x faster than PyTorch** on the forward-plus-backward pass at 579 beams, through the Metal path on Apple Silicon. An optimized CPU path covers macOS and Linux.
|
|
39
39
|
|
|
40
40
|
| Package | Documentation | Source repository | Original project |
|
|
41
41
|
| :--- | :--- | :--- | :--- |
|
|
@@ -69,32 +69,49 @@ the accelerated results against diffBloch's, field by field, including gradients
|
|
|
69
69
|
The diffBloch run makes 4855 calls into the native matrix exponential, so the
|
|
70
70
|
accelerated path is exercised rather than skipped.
|
|
71
71
|
|
|
72
|
+
It also reproduces the published quartz result. Running `diffbloch-fast infer`
|
|
73
|
+
over the 99-rotation Colmey et al. 2026 quartz dataset through the Metal path
|
|
74
|
+
scores every rotation and gives a mean R_obs of 0.0485, against the 0.0486 the
|
|
75
|
+
example documents.
|
|
76
|
+
|
|
72
77
|
---
|
|
73
78
|
|
|
74
79
|
## Performance
|
|
75
80
|
|
|
76
|
-
Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch,
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
|
82
|
-
|
|
|
83
|
-
|
|
|
84
|
-
|
|
|
85
|
-
|
|
|
86
|
-
|
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
81
|
+
Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch, over
|
|
82
|
+
the same diffBloch quartz workload at five beam counts. Both device paths, best
|
|
83
|
+
of three runs each.
|
|
84
|
+
|
|
85
|
+
| Beams | PyTorch | CPU path | Metal path | CPU | Metal |
|
|
86
|
+
| :---: | :---: | :---: | :---: | :---: | :---: |
|
|
87
|
+
| 31 | 1.52 ms | 1.42 ms | 1.95 ms | 1.07x | 0.78x |
|
|
88
|
+
| 61 | 2.11 ms | 2.07 ms | 2.20 ms | 1.02x | 0.96x |
|
|
89
|
+
| 91 | 3.87 ms | 3.61 ms | 3.70 ms | 1.07x | 1.05x |
|
|
90
|
+
| 163 | 10.46 ms | 8.60 ms | 6.27 ms | 1.22x | 1.67x |
|
|
91
|
+
| 579 | 123.30 ms | 89.17 ms | 47.63 ms | 1.38x | **2.59x** |
|
|
92
|
+
|
|
93
|
+
The speedup grows with beam count, and this is the shape of the result rather
|
|
94
|
+
than noise. The package replaces the matrix exponential, which costs O(N^3),
|
|
95
|
+
while structure factors and the loss stay in PyTorch. As the system grows the
|
|
96
|
+
exponential takes a larger share of the runtime, so there is more for the
|
|
97
|
+
accelerated path to reach.
|
|
98
|
+
|
|
99
|
+
That has two practical consequences. At 579 beams, CsPbBr3 scale, Metal runs the
|
|
100
|
+
forward-plus-backward pass in 47.63 ms against PyTorch's 123.30 ms. Below about
|
|
101
|
+
91 beams Metal is slower than PyTorch, because the dispatch cost outweighs a
|
|
102
|
+
small exponential, so `enable(device="cpu")` is the better choice for small
|
|
103
|
+
systems.
|
|
104
|
+
|
|
105
|
+
Forward pass alone at 579 beams: PyTorch 25.90 ms, Metal 15.58 ms, a factor of
|
|
106
|
+
1.66. PyTorch MPS has no native kernel for `aten::linalg_matrix_exp` and falls
|
|
107
|
+
back to the CPU with host transfers, which is the gap this closes.
|
|
91
108
|
|
|
92
109
|
### The standalone Flow port
|
|
93
110
|
|
|
94
111
|
The diffFlow repository also holds a standalone port of the whole calculation,
|
|
95
112
|
compiled from Flow rather than bridged into PyTorch. It avoids the framework
|
|
96
|
-
overhead
|
|
97
|
-
|
|
113
|
+
overhead and goes further. At 579 beams, forward plus backward, minimum of five
|
|
114
|
+
runs:
|
|
98
115
|
|
|
99
116
|
| Implementation | Forward + Backward |
|
|
100
117
|
| :--- | :---: |
|
|
@@ -103,28 +120,12 @@ backward, minimum of five runs:
|
|
|
103
120
|
| Flow port, C backend | 82.84 ms |
|
|
104
121
|
| Flow port, MLIR backend | 79.74 ms |
|
|
105
122
|
|
|
106
|
-
|
|
107
|
-
running the port directly, and they are why the package exists. See the
|
|
123
|
+
Running the port directly is how those are reached. See the
|
|
108
124
|
[diffFlow documentation](https://godofecht.github.io/diffFlow/).
|
|
109
125
|
|
|
110
126
|
These figures come from one workload on one machine. Other crystals, hardware
|
|
111
127
|
configurations and beam counts will differ.
|
|
112
128
|
|
|
113
|
-
### Metal GPU
|
|
114
|
-
|
|
115
|
-
The Metal path is available on Apple Silicon through `enable(device="gpu")`.
|
|
116
|
-
The table below is from an earlier run and has not been reproduced against the
|
|
117
|
-
current package, so treat it as indicative.
|
|
118
|
-
|
|
119
|
-
### Package benchmark snapshot
|
|
120
|
-
|
|
121
|
-
| Implementation | Forward | Forward + Backward | Speedup vs PyTorch CPU | Speedup vs PyTorch MPS |
|
|
122
|
-
| :--- | :---: | :---: | :---: | :---: |
|
|
123
|
-
| PyTorch CPU | 25.7 ms | 130.3 ms | 1.00x | 1.17x |
|
|
124
|
-
| PyTorch MPS (fallback) | 26.2 ms | 153.0 ms | 0.85x | 1.00x |
|
|
125
|
-
| **faster-diffBloch CPU** | **24.4 ms** | **83.5 ms** | **1.56x** | **1.83x** |
|
|
126
|
-
| **faster-diffBloch Metal GPU** | **13.1 ms** | **58.1 ms** | **2.24x** | **2.63x** |
|
|
127
|
-
|
|
128
129
|
---
|
|
129
130
|
|
|
130
131
|
## Installation
|
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Drop-in Apple Silicon Metal GPU and optimized CPU acceleration for [diffBloch](https://diffbloch.com) electron crystallography structure refinement.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
**2.59x faster than PyTorch** on the forward-plus-backward pass at 579 beams, through the Metal path on Apple Silicon. An optimized CPU path covers macOS and Linux.
|
|
6
6
|
|
|
7
7
|
| Package | Documentation | Source repository | Original project |
|
|
8
8
|
| :--- | :--- | :--- | :--- |
|
|
@@ -36,32 +36,49 @@ the accelerated results against diffBloch's, field by field, including gradients
|
|
|
36
36
|
The diffBloch run makes 4855 calls into the native matrix exponential, so the
|
|
37
37
|
accelerated path is exercised rather than skipped.
|
|
38
38
|
|
|
39
|
+
It also reproduces the published quartz result. Running `diffbloch-fast infer`
|
|
40
|
+
over the 99-rotation Colmey et al. 2026 quartz dataset through the Metal path
|
|
41
|
+
scores every rotation and gives a mean R_obs of 0.0485, against the 0.0486 the
|
|
42
|
+
example documents.
|
|
43
|
+
|
|
39
44
|
---
|
|
40
45
|
|
|
41
46
|
## Performance
|
|
42
47
|
|
|
43
|
-
Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch,
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
|
49
|
-
|
|
|
50
|
-
|
|
|
51
|
-
|
|
|
52
|
-
|
|
|
53
|
-
|
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
48
|
+
Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch, over
|
|
49
|
+
the same diffBloch quartz workload at five beam counts. Both device paths, best
|
|
50
|
+
of three runs each.
|
|
51
|
+
|
|
52
|
+
| Beams | PyTorch | CPU path | Metal path | CPU | Metal |
|
|
53
|
+
| :---: | :---: | :---: | :---: | :---: | :---: |
|
|
54
|
+
| 31 | 1.52 ms | 1.42 ms | 1.95 ms | 1.07x | 0.78x |
|
|
55
|
+
| 61 | 2.11 ms | 2.07 ms | 2.20 ms | 1.02x | 0.96x |
|
|
56
|
+
| 91 | 3.87 ms | 3.61 ms | 3.70 ms | 1.07x | 1.05x |
|
|
57
|
+
| 163 | 10.46 ms | 8.60 ms | 6.27 ms | 1.22x | 1.67x |
|
|
58
|
+
| 579 | 123.30 ms | 89.17 ms | 47.63 ms | 1.38x | **2.59x** |
|
|
59
|
+
|
|
60
|
+
The speedup grows with beam count, and this is the shape of the result rather
|
|
61
|
+
than noise. The package replaces the matrix exponential, which costs O(N^3),
|
|
62
|
+
while structure factors and the loss stay in PyTorch. As the system grows the
|
|
63
|
+
exponential takes a larger share of the runtime, so there is more for the
|
|
64
|
+
accelerated path to reach.
|
|
65
|
+
|
|
66
|
+
That has two practical consequences. At 579 beams, CsPbBr3 scale, Metal runs the
|
|
67
|
+
forward-plus-backward pass in 47.63 ms against PyTorch's 123.30 ms. Below about
|
|
68
|
+
91 beams Metal is slower than PyTorch, because the dispatch cost outweighs a
|
|
69
|
+
small exponential, so `enable(device="cpu")` is the better choice for small
|
|
70
|
+
systems.
|
|
71
|
+
|
|
72
|
+
Forward pass alone at 579 beams: PyTorch 25.90 ms, Metal 15.58 ms, a factor of
|
|
73
|
+
1.66. PyTorch MPS has no native kernel for `aten::linalg_matrix_exp` and falls
|
|
74
|
+
back to the CPU with host transfers, which is the gap this closes.
|
|
58
75
|
|
|
59
76
|
### The standalone Flow port
|
|
60
77
|
|
|
61
78
|
The diffFlow repository also holds a standalone port of the whole calculation,
|
|
62
79
|
compiled from Flow rather than bridged into PyTorch. It avoids the framework
|
|
63
|
-
overhead
|
|
64
|
-
|
|
80
|
+
overhead and goes further. At 579 beams, forward plus backward, minimum of five
|
|
81
|
+
runs:
|
|
65
82
|
|
|
66
83
|
| Implementation | Forward + Backward |
|
|
67
84
|
| :--- | :---: |
|
|
@@ -70,28 +87,12 @@ backward, minimum of five runs:
|
|
|
70
87
|
| Flow port, C backend | 82.84 ms |
|
|
71
88
|
| Flow port, MLIR backend | 79.74 ms |
|
|
72
89
|
|
|
73
|
-
|
|
74
|
-
running the port directly, and they are why the package exists. See the
|
|
90
|
+
Running the port directly is how those are reached. See the
|
|
75
91
|
[diffFlow documentation](https://godofecht.github.io/diffFlow/).
|
|
76
92
|
|
|
77
93
|
These figures come from one workload on one machine. Other crystals, hardware
|
|
78
94
|
configurations and beam counts will differ.
|
|
79
95
|
|
|
80
|
-
### Metal GPU
|
|
81
|
-
|
|
82
|
-
The Metal path is available on Apple Silicon through `enable(device="gpu")`.
|
|
83
|
-
The table below is from an earlier run and has not been reproduced against the
|
|
84
|
-
current package, so treat it as indicative.
|
|
85
|
-
|
|
86
|
-
### Package benchmark snapshot
|
|
87
|
-
|
|
88
|
-
| Implementation | Forward | Forward + Backward | Speedup vs PyTorch CPU | Speedup vs PyTorch MPS |
|
|
89
|
-
| :--- | :---: | :---: | :---: | :---: |
|
|
90
|
-
| PyTorch CPU | 25.7 ms | 130.3 ms | 1.00x | 1.17x |
|
|
91
|
-
| PyTorch MPS (fallback) | 26.2 ms | 153.0 ms | 0.85x | 1.00x |
|
|
92
|
-
| **faster-diffBloch CPU** | **24.4 ms** | **83.5 ms** | **1.56x** | **1.83x** |
|
|
93
|
-
| **faster-diffBloch Metal GPU** | **13.1 ms** | **58.1 ms** | **2.24x** | **2.63x** |
|
|
94
|
-
|
|
95
96
|
---
|
|
96
97
|
|
|
97
98
|
## Installation
|
|
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "faster-diffbloch"
|
|
7
|
-
version = "0.1.
|
|
7
|
+
version = "0.1.6"
|
|
8
8
|
description = "Drop-in Metal GPU (macOS) and optimized CPU (macOS/Linux) acceleration for diffBloch"
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
requires-python = ">=3.10"
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
{faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/batch_cgemm.metal
RENAMED
|
File without changes
|
|
File without changes
|
{faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/metal_batch_cgemm.m
RENAMED
|
File without changes
|
{faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/native_scattering.c
RENAMED
|
File without changes
|
{faster_diffbloch-0.1.4 → faster_diffbloch-0.1.6}/src/faster_diffbloch/native/native_scattering.h
RENAMED
|
File without changes
|