DeepGPR 0.0.12__tar.gz → 0.0.13__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,15 +1,19 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: DeepGPR
3
- Version: 0.0.12
3
+ Version: 0.0.13
4
4
  Summary: PyTorch and CUDA for GPR FWI
5
5
  Author-email: Lei Liu <liulei990222@gmail.com>
6
6
  Classifier: Programming Language :: Python :: 3
7
7
  Classifier: Operating System :: Microsoft :: Windows
8
8
  Classifier: Operating System :: POSIX :: Linux
9
+ Classifier: Topic :: Scientific/Engineering
10
+ Classifier: Topic :: Scientific/Engineering :: Physics
9
11
  Requires-Python: >=3.8
10
12
  Description-Content-Type: text/markdown
11
13
  Requires-Dist: numpy
14
+ Requires-Dist: scipy
12
15
  Requires-Dist: matplotlib
16
+ Requires-Dist: torch
13
17
 
14
18
  # DeepGPR
15
19
 
@@ -24,22 +28,26 @@ Gradients of the output receiver data can be computed with respect to model para
24
28
 
25
29
  Utilizes CPML, allowing the width of the PML layer to be configured independently for each boundary.
26
30
 
27
- All operations are executed on the GPU.
31
+ The compute backend can run on CUDA GPUs or on CPU. The CPU backend is implemented in C and is selected automatically when `device='cpu'`.
32
+
33
+ The FDTD spatial finite-difference order can be selected with `fdtd_order=2`, `4`, or `8` (default: `2`).
34
+
35
+ The FWI gradient mode can be selected with `mode=2` or `mode=3`. `mode=2` keeps the previous Ez-only gradient behavior, while `mode=3` uses Ex, Ey, and Ez forward/adjoint electric-field contributions for relative permittivity and conductivity gradients.
28
36
 
29
37
  Supports techniques such as checkpointing, DDP, and the utilization of CPU memory to minimize GPU memory consumption, thereby enabling the execution of large-scale models.
30
38
 
31
39
 
32
40
  ## System Requirements
33
41
 
34
- - **OS**: Linux and Windows
35
- - **Environment**: Python 3.8+, CUDA Toolkit
36
- - **Libraries**: `torch` (with CUDA support), `numpy`, `scipy`, `matplotlib`
37
- - **Hardware**: NVIDIA GPU with sufficient VRAM for 3D computational grids.
42
+ - **OS**: Linux, Windows, and macOS for CPU execution; Linux and Windows for CUDA execution
43
+ - **Environment**: Python 3.8+, CUDA Toolkit for CUDA execution
44
+ - **Libraries**: `torch`, `numpy`, `scipy`, `matplotlib`
45
+ - **Hardware**: NVIDIA GPU with sufficient VRAM for CUDA execution; CPU execution works without a GPU.
38
46
 
39
47
 
40
48
  ## Start
41
49
 
42
- Before use, you must ensure that you have an NVIDIA graphics card and have installed a CUDA-enabled version of PyTorch.
50
+ Before CUDA use, you must ensure that you have an NVIDIA graphics card and have installed a CUDA-enabled version of PyTorch. For CPU use, install a CPU build of PyTorch and include a compiled `deepgpr_cpu` shared library in `src/DeepGPR/lib`.
43
51
 
44
52
  DeepGPR can then be installed using
45
53
 
@@ -57,7 +65,7 @@ import DeepGPR
57
65
  import matplotlib.pyplot as plt
58
66
 
59
67
  # Set up the parameters and models
60
- device=torch.device("cuda")
68
+ device=torch.device("cuda" if torch.cuda.is_available() else "cpu")
61
69
  dx=0.02
62
70
  dt=3e-11
63
71
  nt=2000
@@ -79,7 +87,8 @@ r = DeepGPR.compute(
79
87
  source_amplitudes=source_amplitudes,
80
88
  source_location=source_location,
81
89
  receiver_location=receiver_location,
82
- er=er, se=se
90
+ er=er, se=se,
91
+ fdtd_order=2
83
92
  )
84
93
 
85
94
  (r[-1]**2).sum().backward()
@@ -133,16 +142,20 @@ def compute(device, dx=None, dt=None,
133
142
  E=None, H=None, PML=None,
134
143
  pmlthick=10, source_direction=2, reciever_direction=2,
135
144
  model_gradient_sampling_interval=1,
136
- use_async_offload=False):
145
+ use_async_offload=False,
146
+ fdtd_order=2,
147
+ mode=2):
137
148
  ```
138
149
  ## 📥 Input Parameters
139
150
  ### 1. Basic Physics & Grid Parameters
140
151
 
141
152
  | Parameter | Data Type | Description |
142
153
  | :--- | :--- | :--- |
143
- | **`device`** | `torch.device` / `str` | PyTorch computation device, e.g., `'cuda:0'`. Determines where the computation takes place. |
154
+ | **`device`** | `torch.device` / `str` | PyTorch computation device, e.g., `'cuda:0'` or `'cpu'`. CUDA loads `deepgpr.so/.dll`; CPU loads `deepgpr_cpu.so/.dll/.dylib`. |
144
155
  | **`dx`** | `float` | Spatial grid step size (assuming an isotropic grid, i.e., $dx = dy = dz$). Typically in meters (m). |
145
156
  | **`dt`** | `float` | Time step size. **Note**: Must strictly satisfy the CFL (Courant-Friedrichs-Lewy) stability condition, or an exception will be raised. Typically in seconds (s). |
157
+ | **`fdtd_order`** | `int` | Spatial finite-difference order used by the FDTD field updates. Supported values are `2`, `4`, and `8`; default is `2` for compatibility with earlier versions. |
158
+ | **`mode`** | `int` | FWI gradient mode. `2` keeps the previous Ez-only model-gradient calculation. `3` uses Ex, Ey, and Ez electric-field contributions for relative permittivity and conductivity gradients. |
146
159
  ### 2. Medium Model Parameters
147
160
 
148
161
  This section defines the electromagnetic properties of the simulation space. For 2D simulations, set `nz=1`.
@@ -161,7 +174,7 @@ This section defines the geometric observation system (coordinates) and the exci
161
174
 
162
175
  | Parameter | Data Type | Shape | Description |
163
176
  | :--- | :--- | :--- | :--- |
164
- | **`source_amplitudes`** | `Tensor` (float) | `(num_waveforms, nt)` | Source excitation waveforms. `nt` is the total number of time steps.<br>- If `num_waveforms == 1`: All sources share this single waveform.<br>- If `num_waveforms == nsr`: Each source uses its corresponding waveform. |
177
+ | **`source_amplitudes`** | `Tensor` (float) | `(num_waveforms, nt, 1)` | Source excitation waveforms. `nt` is the total number of time steps.<br>- If `num_waveforms == 1`: All sources share this single waveform.<br>- If `num_waveforms == nsr`: Each source uses its corresponding waveform. |
165
178
  | **`source_location`** | `Tensor` (int) | `(nstep, nsr, 3)` | Grid coordinate indices of the sources.<br>The last dimension corresponds to `[x_idx, y_idx, z_idx]`. |
166
179
  | **`receiver_location`** | `Tensor` (int) | `(nstep, nrx, 3)` | Grid coordinate indices of the receivers.<br>The last dimension corresponds to `[x_idx, y_idx, z_idx]`. |
167
180
  | **`source_direction`** | `int` | Scalar | Polarization direction/component of the source excitation.<br>`0` = X, `1` = Y, `2` = Z (e.g., exciting $E_z$). |
@@ -179,7 +192,34 @@ This section defines the geometric observation system (coordinates) and the exci
179
192
  | :--- | :--- | :--- | :--- |
180
193
  | **`pmlthick`** | `int` / `list` / `Tensor`| Scalar or list of 6 | Thickness (in grid layers) of the PML (Perfectly Matched Layer) absorbing boundaries.<br>- Integer `p`: All six boundaries have thickness `p` (Z-boundaries are ignored in 2D).<br>- List `[x0, xm, y0, ym, z0, zm]`: Specific thicknesses for the 6 boundaries. |
181
194
  | **`model_gradient_sampling_interval`**| `int` | Scalar | Wavefield sampling interval during forward propagation (Default: 1).<br>A larger integer reduces the VRAM usage for the saved `Eall` tensor, but may decrease the accuracy of backpropagated gradients. |
182
- | **`use_async_offload`** | `bool` | Scalar | VRAM optimization flag (Default: `False`).<br>If `True`, the internal full wavefield tensor (`Eall`) is asynchronously offloaded to page-locked host memory (`pin_memory` CPU RAM). This drastically reduces GPU VRAM consumption at the cost of slightly slower computation times due to PCIe data transfer latency. |
195
+ | **`use_async_offload`** | `bool` | Scalar | CUDA-only VRAM optimization flag (Default: `False`).<br>If `True`, the internal full wavefield tensor (`Eall`) is asynchronously offloaded to page-locked host memory (`pin_memory` CPU RAM). This drastically reduces GPU VRAM consumption at the cost of slightly slower computation times due to PCIe data transfer latency. On CPU this option is ignored. |
196
+
197
+ ### 4.1 FWI Gradient Mode
198
+
199
+ `mode` only changes how the model gradients are accumulated during backpropagation:
200
+
201
+ - `mode=2` (default): Saves Ez in `Eall` and computes relative permittivity/conductivity gradients from Ez only. This keeps the old behavior.
202
+ - `mode=3`: Saves Ex, Ey, and Ez in `Eall` and computes relative permittivity/conductivity gradients from all three electric-field components. This is intended for complete 3D Maxwell FWI. The adjoint source polarization is not changed by this option.
203
+
204
+ ## CPU Backend Build
205
+
206
+ The CPU backend is a plain C shared library and is built with OpenMP by default. Build it into `src/DeepGPR/lib` before running with `device='cpu'`. You can control CPU thread count with `OMP_NUM_THREADS`.
207
+
208
+ ```bash
209
+ # Linux
210
+ cc -std=c99 -O3 -fopenmp -fPIC -shared -o src/DeepGPR/lib/deepgpr_cpu.so src/DeepGPR/lib/deepgpr_cpu.c
211
+
212
+ # macOS
213
+ brew install libomp
214
+ LIBOMP_PREFIX="$(brew --prefix libomp)"
215
+ cc -std=c99 -O3 -Xpreprocessor -fopenmp -DDEEPGPR_USE_OPENMP -I"$LIBOMP_PREFIX/include" -Wl,-rpath,"$LIBOMP_PREFIX/lib" -L"$LIBOMP_PREFIX/lib" -fPIC -shared -o src/DeepGPR/lib/deepgpr_cpu.dylib src/DeepGPR/lib/deepgpr_cpu.c -lomp
216
+ ```
217
+
218
+ On Windows, build `src\DeepGPR\lib\deepgpr_cpu.dll` with MSVC:
219
+
220
+ ```powershell
221
+ cl /LD /O2 /openmp /Fe:src\DeepGPR\lib\deepgpr_cpu.dll src\DeepGPR\lib\deepgpr_cpu.c
222
+ ```
183
223
 
184
224
  ### 5. Field Variable States (Checkpoints / Initial Fields)
185
225
 
@@ -201,8 +241,10 @@ The function returns a tuple of 5 elements. These are used to extract synthetic
201
241
  return Eall, (Ex, Ey, Ez), (Hx, Hy, Hz), (x0EPhi1...zmHPhi2), receiver_amplitudes
202
242
  ```
203
243
 
204
- 1. **`Eall`**: The global electric field history saved for gradient calculation.
205
- * **Shape**: `(nt_saved, nstep, nx, ny, nz)` (where `nt_saved` depends on `nt` and `model_gradient_sampling_interval`).
244
+ 1. **`Eall`**: The electric field history saved for gradient calculation.
245
+ * **Shape when `mode=2`**: `(nt_saved, nstep, nx, ny, nz)`, storing Ez only.
246
+ * **Shape when `mode=3`**: `(3, nt_saved, nstep, nx, ny, nz)`, storing components in `[Ex, Ey, Ez]` order.
247
+ * `nt_saved` depends on `nt` and `model_gradient_sampling_interval`.
206
248
  2. **`(Ex, Ey, Ez)`**: The 3D electric field state at the final time step.
207
249
  3. **`(Hx, Hy, Hz)`**: The 3D magnetic field state at the final time step.
208
250
  4. **`(PML_Tuple)`**: A tuple of 24 Tensors recording the final time step state of the PML auxiliary $\Phi$ variables.
@@ -1,16 +1,3 @@
1
- Metadata-Version: 2.4
2
- Name: DeepGPR
3
- Version: 0.0.12
4
- Summary: PyTorch and CUDA for GPR FWI
5
- Author-email: Lei Liu <liulei990222@gmail.com>
6
- Classifier: Programming Language :: Python :: 3
7
- Classifier: Operating System :: Microsoft :: Windows
8
- Classifier: Operating System :: POSIX :: Linux
9
- Requires-Python: >=3.8
10
- Description-Content-Type: text/markdown
11
- Requires-Dist: numpy
12
- Requires-Dist: matplotlib
13
-
14
1
  # DeepGPR
15
2
 
16
3
  DeepGPR provides a wave propagation module for PyTorch, designed for applications such as Ground Penetrating Radar (GPR) imaging and inversion. Its core concepts are derived from Deepwave. You can use it to perform both forward modeling and backpropagation—thereby enabling the simulation of wave propagation to generate synthetic data—as well as for Full Waveform Inversion (FWI). Furthermore, you can integrate this wave propagation functionality into a larger operational pipeline—incorporating various wavelets, loss functions, and other components—to achieve end-to-end forward and reverse propagation, powered by automatic differentiation and our high-performance operators.
@@ -24,22 +11,26 @@ Gradients of the output receiver data can be computed with respect to model para
24
11
 
25
12
  Utilizes CPML, allowing the width of the PML layer to be configured independently for each boundary.
26
13
 
27
- All operations are executed on the GPU.
14
+ The compute backend can run on CUDA GPUs or on CPU. The CPU backend is implemented in C and is selected automatically when `device='cpu'`.
15
+
16
+ The FDTD spatial finite-difference order can be selected with `fdtd_order=2`, `4`, or `8` (default: `2`).
17
+
18
+ The FWI gradient mode can be selected with `mode=2` or `mode=3`. `mode=2` keeps the previous Ez-only gradient behavior, while `mode=3` uses Ex, Ey, and Ez forward/adjoint electric-field contributions for relative permittivity and conductivity gradients.
28
19
 
29
20
  Supports techniques such as checkpointing, DDP, and the utilization of CPU memory to minimize GPU memory consumption, thereby enabling the execution of large-scale models.
30
21
 
31
22
 
32
23
  ## System Requirements
33
24
 
34
- - **OS**: Linux and Windows
35
- - **Environment**: Python 3.8+, CUDA Toolkit
36
- - **Libraries**: `torch` (with CUDA support), `numpy`, `scipy`, `matplotlib`
37
- - **Hardware**: NVIDIA GPU with sufficient VRAM for 3D computational grids.
25
+ - **OS**: Linux, Windows, and macOS for CPU execution; Linux and Windows for CUDA execution
26
+ - **Environment**: Python 3.8+, CUDA Toolkit for CUDA execution
27
+ - **Libraries**: `torch`, `numpy`, `scipy`, `matplotlib`
28
+ - **Hardware**: NVIDIA GPU with sufficient VRAM for CUDA execution; CPU execution works without a GPU.
38
29
 
39
30
 
40
31
  ## Start
41
32
 
42
- Before use, you must ensure that you have an NVIDIA graphics card and have installed a CUDA-enabled version of PyTorch.
33
+ Before CUDA use, you must ensure that you have an NVIDIA graphics card and have installed a CUDA-enabled version of PyTorch. For CPU use, install a CPU build of PyTorch and include a compiled `deepgpr_cpu` shared library in `src/DeepGPR/lib`.
43
34
 
44
35
  DeepGPR can then be installed using
45
36
 
@@ -57,7 +48,7 @@ import DeepGPR
57
48
  import matplotlib.pyplot as plt
58
49
 
59
50
  # Set up the parameters and models
60
- device=torch.device("cuda")
51
+ device=torch.device("cuda" if torch.cuda.is_available() else "cpu")
61
52
  dx=0.02
62
53
  dt=3e-11
63
54
  nt=2000
@@ -79,7 +70,8 @@ r = DeepGPR.compute(
79
70
  source_amplitudes=source_amplitudes,
80
71
  source_location=source_location,
81
72
  receiver_location=receiver_location,
82
- er=er, se=se
73
+ er=er, se=se,
74
+ fdtd_order=2
83
75
  )
84
76
 
85
77
  (r[-1]**2).sum().backward()
@@ -133,16 +125,20 @@ def compute(device, dx=None, dt=None,
133
125
  E=None, H=None, PML=None,
134
126
  pmlthick=10, source_direction=2, reciever_direction=2,
135
127
  model_gradient_sampling_interval=1,
136
- use_async_offload=False):
128
+ use_async_offload=False,
129
+ fdtd_order=2,
130
+ mode=2):
137
131
  ```
138
132
  ## 📥 Input Parameters
139
133
  ### 1. Basic Physics & Grid Parameters
140
134
 
141
135
  | Parameter | Data Type | Description |
142
136
  | :--- | :--- | :--- |
143
- | **`device`** | `torch.device` / `str` | PyTorch computation device, e.g., `'cuda:0'`. Determines where the computation takes place. |
137
+ | **`device`** | `torch.device` / `str` | PyTorch computation device, e.g., `'cuda:0'` or `'cpu'`. CUDA loads `deepgpr.so/.dll`; CPU loads `deepgpr_cpu.so/.dll/.dylib`. |
144
138
  | **`dx`** | `float` | Spatial grid step size (assuming an isotropic grid, i.e., $dx = dy = dz$). Typically in meters (m). |
145
139
  | **`dt`** | `float` | Time step size. **Note**: Must strictly satisfy the CFL (Courant-Friedrichs-Lewy) stability condition, or an exception will be raised. Typically in seconds (s). |
140
+ | **`fdtd_order`** | `int` | Spatial finite-difference order used by the FDTD field updates. Supported values are `2`, `4`, and `8`; default is `2` for compatibility with earlier versions. |
141
+ | **`mode`** | `int` | FWI gradient mode. `2` keeps the previous Ez-only model-gradient calculation. `3` uses Ex, Ey, and Ez electric-field contributions for relative permittivity and conductivity gradients. |
146
142
  ### 2. Medium Model Parameters
147
143
 
148
144
  This section defines the electromagnetic properties of the simulation space. For 2D simulations, set `nz=1`.
@@ -161,7 +157,7 @@ This section defines the geometric observation system (coordinates) and the exci
161
157
 
162
158
  | Parameter | Data Type | Shape | Description |
163
159
  | :--- | :--- | :--- | :--- |
164
- | **`source_amplitudes`** | `Tensor` (float) | `(num_waveforms, nt)` | Source excitation waveforms. `nt` is the total number of time steps.<br>- If `num_waveforms == 1`: All sources share this single waveform.<br>- If `num_waveforms == nsr`: Each source uses its corresponding waveform. |
160
+ | **`source_amplitudes`** | `Tensor` (float) | `(num_waveforms, nt, 1)` | Source excitation waveforms. `nt` is the total number of time steps.<br>- If `num_waveforms == 1`: All sources share this single waveform.<br>- If `num_waveforms == nsr`: Each source uses its corresponding waveform. |
165
161
  | **`source_location`** | `Tensor` (int) | `(nstep, nsr, 3)` | Grid coordinate indices of the sources.<br>The last dimension corresponds to `[x_idx, y_idx, z_idx]`. |
166
162
  | **`receiver_location`** | `Tensor` (int) | `(nstep, nrx, 3)` | Grid coordinate indices of the receivers.<br>The last dimension corresponds to `[x_idx, y_idx, z_idx]`. |
167
163
  | **`source_direction`** | `int` | Scalar | Polarization direction/component of the source excitation.<br>`0` = X, `1` = Y, `2` = Z (e.g., exciting $E_z$). |
@@ -179,7 +175,34 @@ This section defines the geometric observation system (coordinates) and the exci
179
175
  | :--- | :--- | :--- | :--- |
180
176
  | **`pmlthick`** | `int` / `list` / `Tensor`| Scalar or list of 6 | Thickness (in grid layers) of the PML (Perfectly Matched Layer) absorbing boundaries.<br>- Integer `p`: All six boundaries have thickness `p` (Z-boundaries are ignored in 2D).<br>- List `[x0, xm, y0, ym, z0, zm]`: Specific thicknesses for the 6 boundaries. |
181
177
  | **`model_gradient_sampling_interval`**| `int` | Scalar | Wavefield sampling interval during forward propagation (Default: 1).<br>A larger integer reduces the VRAM usage for the saved `Eall` tensor, but may decrease the accuracy of backpropagated gradients. |
182
- | **`use_async_offload`** | `bool` | Scalar | VRAM optimization flag (Default: `False`).<br>If `True`, the internal full wavefield tensor (`Eall`) is asynchronously offloaded to page-locked host memory (`pin_memory` CPU RAM). This drastically reduces GPU VRAM consumption at the cost of slightly slower computation times due to PCIe data transfer latency. |
178
+ | **`use_async_offload`** | `bool` | Scalar | CUDA-only VRAM optimization flag (Default: `False`).<br>If `True`, the internal full wavefield tensor (`Eall`) is asynchronously offloaded to page-locked host memory (`pin_memory` CPU RAM). This drastically reduces GPU VRAM consumption at the cost of slightly slower computation times due to PCIe data transfer latency. On CPU this option is ignored. |
179
+
180
+ ### 4.1 FWI Gradient Mode
181
+
182
+ `mode` only changes how the model gradients are accumulated during backpropagation:
183
+
184
+ - `mode=2` (default): Saves Ez in `Eall` and computes relative permittivity/conductivity gradients from Ez only. This keeps the old behavior.
185
+ - `mode=3`: Saves Ex, Ey, and Ez in `Eall` and computes relative permittivity/conductivity gradients from all three electric-field components. This is intended for complete 3D Maxwell FWI. The adjoint source polarization is not changed by this option.
186
+
187
+ ## CPU Backend Build
188
+
189
+ The CPU backend is a plain C shared library and is built with OpenMP by default. Build it into `src/DeepGPR/lib` before running with `device='cpu'`. You can control CPU thread count with `OMP_NUM_THREADS`.
190
+
191
+ ```bash
192
+ # Linux
193
+ cc -std=c99 -O3 -fopenmp -fPIC -shared -o src/DeepGPR/lib/deepgpr_cpu.so src/DeepGPR/lib/deepgpr_cpu.c
194
+
195
+ # macOS
196
+ brew install libomp
197
+ LIBOMP_PREFIX="$(brew --prefix libomp)"
198
+ cc -std=c99 -O3 -Xpreprocessor -fopenmp -DDEEPGPR_USE_OPENMP -I"$LIBOMP_PREFIX/include" -Wl,-rpath,"$LIBOMP_PREFIX/lib" -L"$LIBOMP_PREFIX/lib" -fPIC -shared -o src/DeepGPR/lib/deepgpr_cpu.dylib src/DeepGPR/lib/deepgpr_cpu.c -lomp
199
+ ```
200
+
201
+ On Windows, build `src\DeepGPR\lib\deepgpr_cpu.dll` with MSVC:
202
+
203
+ ```powershell
204
+ cl /LD /O2 /openmp /Fe:src\DeepGPR\lib\deepgpr_cpu.dll src\DeepGPR\lib\deepgpr_cpu.c
205
+ ```
183
206
 
184
207
  ### 5. Field Variable States (Checkpoints / Initial Fields)
185
208
 
@@ -201,8 +224,10 @@ The function returns a tuple of 5 elements. These are used to extract synthetic
201
224
  return Eall, (Ex, Ey, Ez), (Hx, Hy, Hz), (x0EPhi1...zmHPhi2), receiver_amplitudes
202
225
  ```
203
226
 
204
- 1. **`Eall`**: The global electric field history saved for gradient calculation.
205
- * **Shape**: `(nt_saved, nstep, nx, ny, nz)` (where `nt_saved` depends on `nt` and `model_gradient_sampling_interval`).
227
+ 1. **`Eall`**: The electric field history saved for gradient calculation.
228
+ * **Shape when `mode=2`**: `(nt_saved, nstep, nx, ny, nz)`, storing Ez only.
229
+ * **Shape when `mode=3`**: `(3, nt_saved, nstep, nx, ny, nz)`, storing components in `[Ex, Ey, Ez]` order.
230
+ * `nt_saved` depends on `nt` and `model_gradient_sampling_interval`.
206
231
  2. **`(Ex, Ey, Ez)`**: The 3D electric field state at the final time step.
207
232
  3. **`(Hx, Hy, Hz)`**: The 3D magnetic field state at the final time step.
208
233
  4. **`(PML_Tuple)`**: A tuple of 24 Tensors recording the final time step state of the PML auxiliary $\Phi$ variables.
@@ -1,31 +1,34 @@
1
1
  [build-system]
2
- requires = [
3
- "setuptools>=68",
4
- "wheel"
5
- ]
2
+ requires = ["setuptools>=70", "wheel"]
6
3
  build-backend = "setuptools.build_meta"
7
4
 
8
5
  [project]
9
6
  name = "DeepGPR"
10
- version = "0.0.12"
7
+ version = "0.0.13"
11
8
  authors = [
12
- { name = "Lei Liu", email = "liulei990222@gmail.com" }
9
+ { name = "Lei Liu", email = "liulei990222@gmail.com" }
13
10
  ]
14
11
  description = "PyTorch and CUDA for GPR FWI"
15
12
  readme = "README.md"
16
13
  requires-python = ">=3.8"
17
- dependencies = [
18
- "numpy",
19
- "matplotlib"
20
- ]
14
+
21
15
  classifiers = [
22
16
  "Programming Language :: Python :: 3",
23
17
  "Operating System :: Microsoft :: Windows",
24
- "Operating System :: POSIX :: Linux"
18
+ "Operating System :: POSIX :: Linux",
19
+ "Topic :: Scientific/Engineering",
20
+ "Topic :: Scientific/Engineering :: Physics"
21
+ ]
22
+
23
+ dependencies = [
24
+ "numpy",
25
+ "scipy",
26
+ "matplotlib",
27
+ "torch"
25
28
  ]
26
29
 
27
30
  [tool.setuptools]
28
- package-dir = {"" = "src"}
31
+ package-dir = { "" = "src" }
29
32
  include-package-data = true
30
33
 
31
34
  [tool.setuptools.packages.find]
@@ -33,6 +36,13 @@ where = ["src"]
33
36
 
34
37
  [tool.setuptools.package-data]
35
38
  DeepGPR = [
39
+ "lib/*.cu",
40
+ "lib/*.cuh",
41
+ "lib/*.h",
42
+ "lib/*.hpp",
43
+ "lib/*.cpp",
44
+ "lib/*.c",
36
45
  "lib/*.dll",
37
- "lib/*.so"
46
+ "lib/*.so",
47
+ "lib/*.pyd"
38
48
  ]
@@ -0,0 +1,293 @@
1
+ from __future__ import annotations
2
+
3
+ import ctypes
4
+ import os
5
+ import platform
6
+ from pathlib import Path
7
+
8
+
9
+ _FLOAT_P = ctypes.POINTER(ctypes.c_float)
10
+ _INT_P = ctypes.POINTER(ctypes.c_int)
11
+
12
+ _PACKAGE_DIR = Path(__file__).resolve().parent
13
+ _LIB_DIR = _PACKAGE_DIR / "lib"
14
+ _SYSTEM_NAME = platform.system()
15
+ _LOADED_LIBS: dict[str, ctypes.CDLL] = {}
16
+
17
+
18
+ def _candidate_library_paths(kind: str) -> list[Path]:
19
+ """Return platform-specific native library candidates.
20
+
21
+ Args:
22
+ kind: Backend kind, either "cuda" or "cpu".
23
+ """
24
+ if kind == "cuda":
25
+ if _SYSTEM_NAME == "Windows":
26
+ return [_LIB_DIR / "deepgpr.dll"]
27
+ if _SYSTEM_NAME == "Linux":
28
+ return [_LIB_DIR / "deepgpr.so", _LIB_DIR / "libdeepgpr.so"]
29
+ return []
30
+
31
+ if kind == "cpu":
32
+ if _SYSTEM_NAME == "Windows":
33
+ return [_LIB_DIR / "deepgpr_cpu.dll"]
34
+ if _SYSTEM_NAME == "Darwin":
35
+ return [_LIB_DIR / "libdeepgpr_cpu.dylib", _LIB_DIR / "deepgpr_cpu.dylib"]
36
+ if _SYSTEM_NAME == "Linux":
37
+ return [_LIB_DIR / "deepgpr_cpu.so", _LIB_DIR / "libdeepgpr_cpu.so"]
38
+ return []
39
+
40
+ raise ValueError(f"Unknown DeepGPR library kind: {kind}")
41
+
42
+
43
+ def _available_library_files() -> list[str]:
44
+ """Return native library filenames currently present in the package.
45
+
46
+ Args:
47
+ None.
48
+ """
49
+ if not _LIB_DIR.is_dir():
50
+ return []
51
+ return sorted(p.name for p in _LIB_DIR.iterdir())
52
+
53
+
54
+ def _add_windows_dll_search_paths() -> None:
55
+ """Add Windows DLL search paths for CUDA, conda, and package libraries.
56
+
57
+ Args:
58
+ None.
59
+ """
60
+ if _SYSTEM_NAME != "Windows" or not hasattr(os, "add_dll_directory"):
61
+ return
62
+
63
+ search_dirs: list[Path] = [_LIB_DIR]
64
+
65
+ cuda_path = os.environ.get("CUDA_PATH")
66
+ if cuda_path:
67
+ search_dirs.append(Path(cuda_path) / "bin")
68
+
69
+ for key, value in os.environ.items():
70
+ if key.startswith("CUDA_PATH_V") and value:
71
+ search_dirs.append(Path(value) / "bin")
72
+
73
+ conda_prefix = os.environ.get("CONDA_PREFIX")
74
+ if conda_prefix:
75
+ search_dirs.append(Path(conda_prefix) / "Library" / "bin")
76
+
77
+ seen = set()
78
+ for directory in search_dirs:
79
+ directory = Path(directory)
80
+ if directory in seen:
81
+ continue
82
+ seen.add(directory)
83
+
84
+ try:
85
+ if directory.is_dir():
86
+ os.add_dll_directory(str(directory))
87
+ except OSError:
88
+ pass
89
+
90
+
91
+ def _require_exported_symbols(lib: ctypes.CDLL, symbols: tuple[str, ...], path: Path) -> None:
92
+ """Validate that a loaded native library exports required symbols.
93
+
94
+ Args:
95
+ lib: Loaded ctypes library object.
96
+ symbols: Symbol names that must be exported.
97
+ path: Filesystem path of the loaded library.
98
+ """
99
+ missing = [name for name in symbols if not hasattr(lib, name)]
100
+ if missing:
101
+ raise RuntimeError(
102
+ "The shared library was loaded, but the following exported C ABI symbols "
103
+ f"were not found: {missing}.\nLibrary: {path}"
104
+ )
105
+
106
+
107
+ def _configure_deepgpr_library(lib: ctypes.CDLL) -> None:
108
+ """Configure ctypes signatures for the native DeepGPR library.
109
+
110
+ Args:
111
+ lib: Loaded ctypes library object to configure.
112
+ """
113
+ if getattr(lib, "_deepgpr_argtypes_configured", False):
114
+ return
115
+
116
+ lib.forward.argtypes = [
117
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
118
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
119
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
120
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
121
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
122
+
123
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
124
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
125
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
126
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
127
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
128
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
129
+
130
+ ctypes.c_int, ctypes.c_int, ctypes.c_int,
131
+ ctypes.c_int, ctypes.c_int, ctypes.c_int,
132
+
133
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
134
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
135
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
136
+
137
+ ctypes.c_float, ctypes.c_int, ctypes.c_int, ctypes.c_int, ctypes.c_float,
138
+ _INT_P, _FLOAT_P,
139
+ ctypes.c_int, ctypes.c_int, ctypes.c_int, ctypes.c_int,
140
+ _INT_P, _FLOAT_P,
141
+ ctypes.c_int, ctypes.c_int, ctypes.c_int,
142
+ ]
143
+ lib.forward.restype = None
144
+
145
+ lib.backward.argtypes = [
146
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
147
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
148
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
149
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
150
+ _FLOAT_P, _FLOAT_P, _FLOAT_P,
151
+
152
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
153
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
154
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
155
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
156
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
157
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
158
+
159
+ ctypes.c_int, ctypes.c_int, ctypes.c_int,
160
+ ctypes.c_int, ctypes.c_int, ctypes.c_int,
161
+
162
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
163
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
164
+ _FLOAT_P, _FLOAT_P, _FLOAT_P, _FLOAT_P,
165
+
166
+ ctypes.c_float, ctypes.c_int, ctypes.c_int, ctypes.c_int, ctypes.c_float,
167
+ ctypes.c_int, ctypes.c_int, ctypes.c_int, ctypes.c_int,
168
+ _INT_P, _FLOAT_P,
169
+ ctypes.c_int,
170
+ _FLOAT_P, _FLOAT_P,
171
+ ctypes.c_int, ctypes.c_int, ctypes.c_int, ctypes.c_int,
172
+ ]
173
+ lib.backward.restype = None
174
+
175
+ if hasattr(lib, "set_fdtd_order"):
176
+ lib.set_fdtd_order.argtypes = [ctypes.c_int]
177
+ lib.set_fdtd_order.restype = None
178
+
179
+ lib._deepgpr_argtypes_configured = True
180
+
181
+
182
+ def _load_deepgpr_library(kind: str) -> ctypes.CDLL:
183
+ """Load and configure a native DeepGPR backend library.
184
+
185
+ Args:
186
+ kind: Backend kind, either "cuda" or "cpu".
187
+ """
188
+ if kind in _LOADED_LIBS:
189
+ return _LOADED_LIBS[kind]
190
+
191
+ _add_windows_dll_search_paths()
192
+ candidates = _candidate_library_paths(kind)
193
+ load_errors: list[str] = []
194
+
195
+ for path in candidates:
196
+ if not path.is_file():
197
+ continue
198
+ try:
199
+ lib = ctypes.CDLL(str(path.resolve()))
200
+ except OSError as exc:
201
+ load_errors.append(f"{path}: {exc}")
202
+ continue
203
+
204
+ _require_exported_symbols(lib, ("forward", "backward"), path)
205
+ _configure_deepgpr_library(lib)
206
+ lib._deepgpr_path = str(path.resolve())
207
+ _LOADED_LIBS[kind] = lib
208
+ return lib
209
+
210
+ expected = "\n".join(str(p) for p in candidates) or "No candidates for this platform."
211
+ details = "\n".join(load_errors) if load_errors else "No load attempts succeeded."
212
+ raise FileNotFoundError(
213
+ f"DeepGPR {kind.upper()} shared library was not found or could not be loaded.\n\n"
214
+ f"Current platform: {_SYSTEM_NAME}\n"
215
+ f"Expected one of:\n{expected}\n\n"
216
+ f"Available files in {_LIB_DIR}:\n{_available_library_files()}\n\n"
217
+ f"Load details:\n{details}"
218
+ )
219
+
220
+
221
+ def _device_type(device) -> str:
222
+ """Convert a PyTorch device or device string to a backend name.
223
+
224
+ Args:
225
+ device: PyTorch device object or device string.
226
+ """
227
+ device_type = getattr(device, "type", None)
228
+ if device_type is not None:
229
+ return str(device_type).lower()
230
+ return str(device).split(":", 1)[0].lower()
231
+
232
+
233
+ def get_deepgpr_lib(device) -> ctypes.CDLL:
234
+ """Return the native library for the requested device.
235
+
236
+ Args:
237
+ device: PyTorch device object or device string.
238
+ """
239
+ device_type = _device_type(device)
240
+ if device_type == "cpu":
241
+ return _load_deepgpr_library("cpu")
242
+ if device_type == "cuda":
243
+ return _load_deepgpr_library("cuda")
244
+ raise ValueError(f"Unsupported DeepGPR device: {device}")
245
+
246
+
247
+ def get_deepgpr_library_path(device) -> str:
248
+ """Return the loaded native library path for a device.
249
+
250
+ Args:
251
+ device: PyTorch device object or device string.
252
+ """
253
+ return str(getattr(get_deepgpr_lib(device), "_deepgpr_path", ""))
254
+
255
+
256
+ def set_library_fdtd_order(lib: ctypes.CDLL, fdtd_order: int) -> None:
257
+ """Set the FDTD spatial order on a native library.
258
+
259
+ Args:
260
+ lib: Loaded ctypes library object.
261
+ fdtd_order: Spatial finite-difference order, supported values are 2, 4, and 8.
262
+ """
263
+ if fdtd_order not in (2, 4, 8):
264
+ raise ValueError("fdtd_order must be one of 2, 4, or 8.")
265
+
266
+ if hasattr(lib, "set_fdtd_order"):
267
+ lib.set_fdtd_order(int(fdtd_order))
268
+ return
269
+
270
+ if fdtd_order != 2:
271
+ raise RuntimeError(
272
+ "The loaded DeepGPR shared library does not support fdtd_order. "
273
+ "Rebuild it from the updated C/CUDA sources to use 4th or 8th order FDTD."
274
+ )
275
+
276
+
277
+ from .common import *
278
+ from .compute2 import *
279
+ from .multiscale import *
280
+ from .wavelet import *
281
+
282
+
283
+ _EXCLUDED_FROM_ALL = {
284
+ "ctypes",
285
+ "os",
286
+ "platform",
287
+ "Path",
288
+ }
289
+
290
+ __all__ = [
291
+ name for name in globals()
292
+ if not name.startswith("_") and name not in _EXCLUDED_FROM_ALL
293
+ ]