faster-diffbloch 0.1.4__tar.gz → 0.1.6__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: faster-diffbloch
3
- Version: 0.1.4
3
+ Version: 0.1.6
4
4
  Summary: Drop-in Metal GPU (macOS) and optimized CPU (macOS/Linux) acceleration for diffBloch
5
5
  Project-URL: Homepage, https://godofecht.github.io/diffFlow/
6
6
  Project-URL: Documentation, https://godofecht.github.io/diffFlow/
@@ -35,7 +35,7 @@ Description-Content-Type: text/markdown
35
35
 
36
36
  Drop-in Apple Silicon Metal GPU and optimized CPU acceleration for [diffBloch](https://diffbloch.com) electron crystallography structure refinement.
37
37
 
38
- A Metal GPU path for Apple Silicon and an optimized CPU path for macOS and Linux, measured against diffBloch running on PyTorch.
38
+ **2.59x faster than PyTorch** on the forward-plus-backward pass at 579 beams, through the Metal path on Apple Silicon. An optimized CPU path covers macOS and Linux.
39
39
 
40
40
  | Package | Documentation | Source repository | Original project |
41
41
  | :--- | :--- | :--- | :--- |
@@ -69,32 +69,49 @@ the accelerated results against diffBloch's, field by field, including gradients
69
69
  The diffBloch run makes 4855 calls into the native matrix exponential, so the
70
70
  accelerated path is exercised rather than skipped.
71
71
 
72
+ It also reproduces the published quartz result. Running `diffbloch-fast infer`
73
+ over the 99-rotation Colmey et al. 2026 quartz dataset through the Metal path
74
+ scores every rotation and gives a mean R_obs of 0.0485, against the 0.0486 the
75
+ example documents.
76
+
72
77
  ---
73
78
 
74
79
  ## Performance
75
80
 
76
- Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch,
77
- running the same diffBloch workload through `bench/compare_faster.py`. This is
78
- what `pip install faster-diffbloch` gives you: the accelerated matrix
79
- exponential bridged into diffBloch, with PyTorch handling everything else.
80
-
81
- | Beams | faster-diffbloch | PyTorch | Speedup |
82
- | :---: | :---: | :---: | :---: |
83
- | 31 | 1.48 ms | 1.54 ms | 1.05x |
84
- | 61 | 1.93 ms | 2.07 ms | 1.07x |
85
- | 91 | 3.48 ms | 3.69 ms | 1.06x |
86
- | 163 | 8.76 ms | 10.14 ms | 1.16x |
87
-
88
- Best of three alternating runs per mode. The gain grows with beam count,
89
- because the matrix exponential takes a larger share of the work as the system
90
- grows and the fixed ctypes and PyTorch overhead matters less.
81
+ Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch, over
82
+ the same diffBloch quartz workload at five beam counts. Both device paths, best
83
+ of three runs each.
84
+
85
+ | Beams | PyTorch | CPU path | Metal path | CPU | Metal |
86
+ | :---: | :---: | :---: | :---: | :---: | :---: |
87
+ | 31 | 1.52 ms | 1.42 ms | 1.95 ms | 1.07x | 0.78x |
88
+ | 61 | 2.11 ms | 2.07 ms | 2.20 ms | 1.02x | 0.96x |
89
+ | 91 | 3.87 ms | 3.61 ms | 3.70 ms | 1.07x | 1.05x |
90
+ | 163 | 10.46 ms | 8.60 ms | 6.27 ms | 1.22x | 1.67x |
91
+ | 579 | 123.30 ms | 89.17 ms | 47.63 ms | 1.38x | **2.59x** |
92
+
93
+ The speedup grows with beam count, and this is the shape of the result rather
94
+ than noise. The package replaces the matrix exponential, which costs O(N^3),
95
+ while structure factors and the loss stay in PyTorch. As the system grows the
96
+ exponential takes a larger share of the runtime, so there is more for the
97
+ accelerated path to reach.
98
+
99
+ That has two practical consequences. At 579 beams, CsPbBr3 scale, Metal runs the
100
+ forward-plus-backward pass in 47.63 ms against PyTorch's 123.30 ms. Below about
101
+ 91 beams Metal is slower than PyTorch, because the dispatch cost outweighs a
102
+ small exponential, so `enable(device="cpu")` is the better choice for small
103
+ systems.
104
+
105
+ Forward pass alone at 579 beams: PyTorch 25.90 ms, Metal 15.58 ms, a factor of
106
+ 1.66. PyTorch MPS has no native kernel for `aten::linalg_matrix_exp` and falls
107
+ back to the CPU with host transfers, which is the gap this closes.
91
108
 
92
109
  ### The standalone Flow port
93
110
 
94
111
  The diffFlow repository also holds a standalone port of the whole calculation,
95
112
  compiled from Flow rather than bridged into PyTorch. It avoids the framework
96
- overhead entirely and is considerably faster. At 579 beams, forward plus
97
- backward, minimum of five runs:
113
+ overhead and goes further. At 579 beams, forward plus backward, minimum of five
114
+ runs:
98
115
 
99
116
  | Implementation | Forward + Backward |
100
117
  | :--- | :---: |
@@ -103,28 +120,12 @@ backward, minimum of five runs:
103
120
  | Flow port, C backend | 82.84 ms |
104
121
  | Flow port, MLIR backend | 79.74 ms |
105
122
 
106
- Those numbers are not what this package delivers. They are reachable by
107
- running the port directly, and they are why the package exists. See the
123
+ Running the port directly is how those are reached. See the
108
124
  [diffFlow documentation](https://godofecht.github.io/diffFlow/).
109
125
 
110
126
  These figures come from one workload on one machine. Other crystals, hardware
111
127
  configurations and beam counts will differ.
112
128
 
113
- ### Metal GPU
114
-
115
- The Metal path is available on Apple Silicon through `enable(device="gpu")`.
116
- The table below is from an earlier run and has not been reproduced against the
117
- current package, so treat it as indicative.
118
-
119
- ### Package benchmark snapshot
120
-
121
- | Implementation | Forward | Forward + Backward | Speedup vs PyTorch CPU | Speedup vs PyTorch MPS |
122
- | :--- | :---: | :---: | :---: | :---: |
123
- | PyTorch CPU | 25.7 ms | 130.3 ms | 1.00x | 1.17x |
124
- | PyTorch MPS (fallback) | 26.2 ms | 153.0 ms | 0.85x | 1.00x |
125
- | **faster-diffBloch CPU** | **24.4 ms** | **83.5 ms** | **1.56x** | **1.83x** |
126
- | **faster-diffBloch Metal GPU** | **13.1 ms** | **58.1 ms** | **2.24x** | **2.63x** |
127
-
128
129
  ---
129
130
 
130
131
  ## Installation
@@ -2,7 +2,7 @@
2
2
 
3
3
  Drop-in Apple Silicon Metal GPU and optimized CPU acceleration for [diffBloch](https://diffbloch.com) electron crystallography structure refinement.
4
4
 
5
- A Metal GPU path for Apple Silicon and an optimized CPU path for macOS and Linux, measured against diffBloch running on PyTorch.
5
+ **2.59x faster than PyTorch** on the forward-plus-backward pass at 579 beams, through the Metal path on Apple Silicon. An optimized CPU path covers macOS and Linux.
6
6
 
7
7
  | Package | Documentation | Source repository | Original project |
8
8
  | :--- | :--- | :--- | :--- |
@@ -36,32 +36,49 @@ the accelerated results against diffBloch's, field by field, including gradients
36
36
  The diffBloch run makes 4855 calls into the native matrix exponential, so the
37
37
  accelerated path is exercised rather than skipped.
38
38
 
39
+ It also reproduces the published quartz result. Running `diffbloch-fast infer`
40
+ over the 99-rotation Colmey et al. 2026 quartz dataset through the Metal path
41
+ scores every rotation and gives a mean R_obs of 0.0485, against the 0.0486 the
42
+ example documents.
43
+
39
44
  ---
40
45
 
41
46
  ## Performance
42
47
 
43
- Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch,
44
- running the same diffBloch workload through `bench/compare_faster.py`. This is
45
- what `pip install faster-diffbloch` gives you: the accelerated matrix
46
- exponential bridged into diffBloch, with PyTorch handling everything else.
47
-
48
- | Beams | faster-diffbloch | PyTorch | Speedup |
49
- | :---: | :---: | :---: | :---: |
50
- | 31 | 1.48 ms | 1.54 ms | 1.05x |
51
- | 61 | 1.93 ms | 2.07 ms | 1.07x |
52
- | 91 | 3.48 ms | 3.69 ms | 1.06x |
53
- | 163 | 8.76 ms | 10.14 ms | 1.16x |
54
-
55
- Best of three alternating runs per mode. The gain grows with beam count,
56
- because the matrix exponential takes a larger share of the work as the system
57
- grows and the fixed ctypes and PyTorch overhead matters less.
48
+ Forward-plus-backward timing on an Apple M4 Max, single-threaded PyTorch, over
49
+ the same diffBloch quartz workload at five beam counts. Both device paths, best
50
+ of three runs each.
51
+
52
+ | Beams | PyTorch | CPU path | Metal path | CPU | Metal |
53
+ | :---: | :---: | :---: | :---: | :---: | :---: |
54
+ | 31 | 1.52 ms | 1.42 ms | 1.95 ms | 1.07x | 0.78x |
55
+ | 61 | 2.11 ms | 2.07 ms | 2.20 ms | 1.02x | 0.96x |
56
+ | 91 | 3.87 ms | 3.61 ms | 3.70 ms | 1.07x | 1.05x |
57
+ | 163 | 10.46 ms | 8.60 ms | 6.27 ms | 1.22x | 1.67x |
58
+ | 579 | 123.30 ms | 89.17 ms | 47.63 ms | 1.38x | **2.59x** |
59
+
60
+ The speedup grows with beam count, and this is the shape of the result rather
61
+ than noise. The package replaces the matrix exponential, which costs O(N^3),
62
+ while structure factors and the loss stay in PyTorch. As the system grows the
63
+ exponential takes a larger share of the runtime, so there is more for the
64
+ accelerated path to reach.
65
+
66
+ That has two practical consequences. At 579 beams, CsPbBr3 scale, Metal runs the
67
+ forward-plus-backward pass in 47.63 ms against PyTorch's 123.30 ms. Below about
68
+ 91 beams Metal is slower than PyTorch, because the dispatch cost outweighs a
69
+ small exponential, so `enable(device="cpu")` is the better choice for small
70
+ systems.
71
+
72
+ Forward pass alone at 579 beams: PyTorch 25.90 ms, Metal 15.58 ms, a factor of
73
+ 1.66. PyTorch MPS has no native kernel for `aten::linalg_matrix_exp` and falls
74
+ back to the CPU with host transfers, which is the gap this closes.
58
75
 
59
76
  ### The standalone Flow port
60
77
 
61
78
  The diffFlow repository also holds a standalone port of the whole calculation,
62
79
  compiled from Flow rather than bridged into PyTorch. It avoids the framework
63
- overhead entirely and is considerably faster. At 579 beams, forward plus
64
- backward, minimum of five runs:
80
+ overhead and goes further. At 579 beams, forward plus backward, minimum of five
81
+ runs:
65
82
 
66
83
  | Implementation | Forward + Backward |
67
84
  | :--- | :---: |
@@ -70,28 +87,12 @@ backward, minimum of five runs:
70
87
  | Flow port, C backend | 82.84 ms |
71
88
  | Flow port, MLIR backend | 79.74 ms |
72
89
 
73
- Those numbers are not what this package delivers. They are reachable by
74
- running the port directly, and they are why the package exists. See the
90
+ Running the port directly is how those are reached. See the
75
91
  [diffFlow documentation](https://godofecht.github.io/diffFlow/).
76
92
 
77
93
  These figures come from one workload on one machine. Other crystals, hardware
78
94
  configurations and beam counts will differ.
79
95
 
80
- ### Metal GPU
81
-
82
- The Metal path is available on Apple Silicon through `enable(device="gpu")`.
83
- The table below is from an earlier run and has not been reproduced against the
84
- current package, so treat it as indicative.
85
-
86
- ### Package benchmark snapshot
87
-
88
- | Implementation | Forward | Forward + Backward | Speedup vs PyTorch CPU | Speedup vs PyTorch MPS |
89
- | :--- | :---: | :---: | :---: | :---: |
90
- | PyTorch CPU | 25.7 ms | 130.3 ms | 1.00x | 1.17x |
91
- | PyTorch MPS (fallback) | 26.2 ms | 153.0 ms | 0.85x | 1.00x |
92
- | **faster-diffBloch CPU** | **24.4 ms** | **83.5 ms** | **1.56x** | **1.83x** |
93
- | **faster-diffBloch Metal GPU** | **13.1 ms** | **58.1 ms** | **2.24x** | **2.63x** |
94
-
95
96
  ---
96
97
 
97
98
  ## Installation
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "faster-diffbloch"
7
- version = "0.1.4"
7
+ version = "0.1.6"
8
8
  description = "Drop-in Metal GPU (macOS) and optimized CPU (macOS/Linux) acceleration for diffBloch"
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.10"
@@ -4,7 +4,7 @@ from __future__ import annotations
4
4
 
5
5
  from .backend import enable, disable, matrix_exp, matrix_exp_backward, faster_propagate
6
6
 
7
- __version__ = "0.1.4"
7
+ __version__ = "0.1.6"
8
8
  __all__ = [
9
9
  "enable",
10
10
  "disable",