cunumpy 0.3.0__tar.gz → 0.6.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (61) hide show
  1. cunumpy-0.6.0/PKG-INFO +575 -0
  2. cunumpy-0.6.0/README.md +535 -0
  3. {cunumpy-0.3.0 → cunumpy-0.6.0}/pyproject.toml +14 -7
  4. cunumpy-0.6.0/src/cunumpy/LLM_GUIDE.md +532 -0
  5. cunumpy-0.6.0/src/cunumpy/__init__.py +169 -0
  6. cunumpy-0.6.0/src/cunumpy/__init__.pyi +47 -0
  7. cunumpy-0.6.0/src/cunumpy/_algorithms.py +253 -0
  8. cunumpy-0.6.0/src/cunumpy/_cuda_kernel.py +2631 -0
  9. cunumpy-0.6.0/src/cunumpy/_device.py +260 -0
  10. cunumpy-0.6.0/src/cunumpy/_dispatch.py +899 -0
  11. cunumpy-0.6.0/src/cunumpy/_emulation.py +446 -0
  12. cunumpy-0.6.0/src/cunumpy/_fake_cupy.py +97 -0
  13. cunumpy-0.6.0/src/cunumpy/_fake_cupy_impl.py +558 -0
  14. cunumpy-0.6.0/src/cunumpy/_fusion.py +115 -0
  15. cunumpy-0.6.0/src/cunumpy/_kernel.py +785 -0
  16. cunumpy-0.6.0/src/cunumpy/_mirror.py +252 -0
  17. cunumpy-0.6.0/src/cunumpy/_morton.py +189 -0
  18. cunumpy-0.6.0/src/cunumpy/_mpi.py +383 -0
  19. cunumpy-0.6.0/src/cunumpy/_philox.py +149 -0
  20. cunumpy-0.6.0/src/cunumpy/_profiling.py +148 -0
  21. cunumpy-0.6.0/src/cunumpy/_random_streams.py +214 -0
  22. cunumpy-0.6.0/src/cunumpy/_scipy_backend.py +157 -0
  23. cunumpy-0.6.0/src/cunumpy/_staging.py +292 -0
  24. cunumpy-0.6.0/src/cunumpy/_streams.py +97 -0
  25. cunumpy-0.6.0/src/cunumpy/_transfers.py +299 -0
  26. cunumpy-0.6.0/src/cunumpy/algorithms.py +39 -0
  27. cunumpy-0.6.0/src/cunumpy/arguments.py +42 -0
  28. cunumpy-0.6.0/src/cunumpy/cuda/__init__.py +90 -0
  29. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/array_view.cuh +240 -0
  30. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/atomic.cuh +110 -0
  31. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/index.cuh +52 -0
  32. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/morton.cuh +128 -0
  33. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/random.cuh +127 -0
  34. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/reduce.cuh +192 -0
  35. cunumpy-0.6.0/src/cunumpy/cuda/include/cunumpy/scan.cuh +84 -0
  36. cunumpy-0.6.0/src/cunumpy/kernel_testing.py +637 -0
  37. cunumpy-0.6.0/src/cunumpy/kernels.py +64 -0
  38. cunumpy-0.6.0/src/cunumpy/memory.py +22 -0
  39. cunumpy-0.6.0/src/cunumpy/mpi.py +81 -0
  40. cunumpy-0.6.0/src/cunumpy/petsc.py +122 -0
  41. cunumpy-0.6.0/src/cunumpy/profiling.py +32 -0
  42. cunumpy-0.6.0/src/cunumpy/rng.py +43 -0
  43. cunumpy-0.6.0/src/cunumpy/xp.py +494 -0
  44. cunumpy-0.6.0/src/cunumpy.egg-info/PKG-INFO +575 -0
  45. cunumpy-0.6.0/src/cunumpy.egg-info/SOURCES.txt +48 -0
  46. {cunumpy-0.3.0 → cunumpy-0.6.0}/src/cunumpy.egg-info/requires.txt +3 -2
  47. cunumpy-0.3.0/PKG-INFO +0 -356
  48. cunumpy-0.3.0/README.md +0 -317
  49. cunumpy-0.3.0/src/cunumpy/__init__.py +0 -102
  50. cunumpy-0.3.0/src/cunumpy/__init__.pyi +0 -55
  51. cunumpy-0.3.0/src/cunumpy/cuda_kernel.py +0 -1003
  52. cunumpy-0.3.0/src/cunumpy/dispatch.py +0 -352
  53. cunumpy-0.3.0/src/cunumpy/kernel.py +0 -360
  54. cunumpy-0.3.0/src/cunumpy/main.py +0 -8
  55. cunumpy-0.3.0/src/cunumpy/xp.py +0 -486
  56. cunumpy-0.3.0/src/cunumpy.egg-info/PKG-INFO +0 -356
  57. cunumpy-0.3.0/src/cunumpy.egg-info/SOURCES.txt +0 -15
  58. {cunumpy-0.3.0 → cunumpy-0.6.0}/setup.cfg +0 -0
  59. {cunumpy-0.3.0 → cunumpy-0.6.0}/src/cunumpy/py.typed +0 -0
  60. {cunumpy-0.3.0 → cunumpy-0.6.0}/src/cunumpy.egg-info/dependency_links.txt +0 -0
  61. {cunumpy-0.3.0 → cunumpy-0.6.0}/src/cunumpy.egg-info/top_level.txt +0 -0
cunumpy-0.6.0/PKG-INFO ADDED
@@ -0,0 +1,575 @@
1
+ Metadata-Version: 2.4
2
+ Name: cunumpy
3
+ Version: 0.6.0
4
+ Summary: Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`.
5
+ Author: Max
6
+ Project-URL: Source, https://github.com/max-models/cunumpy
7
+ Keywords: python
8
+ Classifier: Development Status :: 3 - Alpha
9
+ Classifier: Programming Language :: Python :: 3 :: Only
10
+ Classifier: Programming Language :: Python :: 3.10
11
+ Classifier: Programming Language :: Python :: 3.11
12
+ Classifier: Programming Language :: Python :: 3.12
13
+ Classifier: Programming Language :: Python :: 3.13
14
+ Classifier: Programming Language :: Python :: 3.14
15
+ Requires-Python: >=3.10
16
+ Description-Content-Type: text/markdown
17
+ Requires-Dist: array-api-compat
18
+ Requires-Dist: maybempi>=0.1.2
19
+ Requires-Dist: numpy
20
+ Provides-Extra: dev
21
+ Requires-Dist: ruff; extra == "dev"
22
+ Requires-Dist: cunumpy[docs,test-compiled]; extra == "dev"
23
+ Provides-Extra: docs
24
+ Requires-Dist: ipykernel; extra == "docs"
25
+ Requires-Dist: myst-parser; extra == "docs"
26
+ Requires-Dist: nbconvert; extra == "docs"
27
+ Requires-Dist: nbsphinx; extra == "docs"
28
+ Requires-Dist: jupyterlab; extra == "docs"
29
+ Requires-Dist: pre-commit; extra == "docs"
30
+ Requires-Dist: pyproject-fmt; extra == "docs"
31
+ Requires-Dist: sphinx; extra == "docs"
32
+ Requires-Dist: sphinx-book-theme; extra == "docs"
33
+ Provides-Extra: test
34
+ Requires-Dist: coverage; extra == "test"
35
+ Requires-Dist: pytest; extra == "test"
36
+ Requires-Dist: scipy; extra == "test"
37
+ Provides-Extra: test-compiled
38
+ Requires-Dist: cunumpy[test]; extra == "test-compiled"
39
+ Requires-Dist: pyccel; extra == "test-compiled"
40
+
41
+ # CuNumpy
42
+
43
+ CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays
44
+ on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace
45
+ `import numpy as np` with `import cunumpy as xp`; the array operations you
46
+ already know then run on the selected backend.
47
+
48
+ ```python
49
+ import cunumpy as xp
50
+
51
+ values = xp.arange(5, dtype=xp.float64)
52
+ print(values * 2)
53
+ print(xp.get_backend()) # 'numpy' by default
54
+ ```
55
+
56
+ CuNumpy selects an array library for newly requested operations. It does not
57
+ move existing arrays just because the selected backend changes. This guide
58
+ covers backend selection, array movement, mixed CPU/GPU workflows, and the
59
+ helper APIs CuNumpy provides around NumPy and CuPy.
60
+
61
+ The top level of `cunumpy` is the NumPy (or CuPy) namespace plus backend
62
+ selection and array conversion. The helpers are in submodules, so that they
63
+ never hide a NumPy name:
64
+
65
+ | Submodule | Contents |
66
+ |---|---|
67
+ | `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, host implementations, `fuse` |
68
+ | `xp.arguments` | CUDA only: `CudaStruct`, `CudaStructArguments`, `CudaArguments` |
69
+ | `xp.cuda` | CUDA only: devices, streams, debug mode, CUDA headers |
70
+ | `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
71
+ | `xp.algorithms` | `morton_*`, `sort_by_key`, `cell_offsets`, `segment_boundaries`, `segment_sum`, `SegmentPlan` |
72
+ | `xp.mpi` | `mpi_buffer`, reusable `MPIStaging`, CUDA-aware MPI |
73
+ | `xp.profiling` | `timed_region`, `nvtx_range`, `count_transfers` |
74
+ | `xp.memory` | `HostStaging`, `DeviceMirror` |
75
+ | `xp.petsc` | `petsc_vec` |
76
+ | `cunumpy.kernel_testing` | pytest helpers for host/CUDA kernel pairs |
77
+
78
+ Everything except `xp.cuda`, `xp.arguments` and `CudaKernel` works on both backends.
79
+
80
+ ## Install
81
+
82
+ ```bash
83
+ python -m pip install cunumpy
84
+ ```
85
+
86
+ NumPy and `array-api-compat` are installed as dependencies. To use a GPU,
87
+ install a CuPy package compatible with your CUDA environment as well. CuPy
88
+ installation depends on the CUDA version and platform; follow the CuPy
89
+ installation instructions for your system. CuNumpy does not install CUDA.
90
+
91
+ `array-api-compat` supplies NumPy and CuPy compatibility modules with more
92
+ consistent behavior for shared array operations. CuNumpy uses them internally;
93
+ your arrays remain ordinary NumPy or CuPy arrays. See [why CuNumpy uses
94
+ `array-api-compat`](docs/source/array-api-compat.md) for a plain-language
95
+ explanation and examples.
96
+
97
+ ## Choose a backend
98
+
99
+ CuNumpy starts with NumPy unless `CUNUMPY_BACKEND=cupy` is set before import.
100
+ You can also choose at runtime:
101
+
102
+ ```python
103
+ import cunumpy as xp
104
+
105
+ xp.set_backend("cupy")
106
+ print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional
107
+
108
+ values = xp.arange(5) # created by the active backend
109
+ ```
110
+
111
+ The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested
112
+ but unavailable or not functional, CuNumpy falls back to NumPy. Always check
113
+ `get_backend()` when the effective backend matters, such as when reporting
114
+ configuration or deciding whether GPU-specific work will happen.
115
+
116
+ Use `xp.set_backend("cupy", strict=True)` to raise when CUDA is unavailable,
117
+ preserving the previous backend. `xp.backend_info()` returns structured backend,
118
+ dependency, and CUDA diagnostics. Reusable streams/events, MPI staging, cell
119
+ ranges, and prepared reductions are described in the
120
+ [execution helpers guide](docs/source/guides/execution-helpers.md).
121
+
122
+ Use `use_backend()` for a temporary selection. It restores the previous
123
+ selection when the block exits, including when an exception is raised:
124
+
125
+ ```python
126
+ with xp.use_backend("numpy"):
127
+ cpu_values = xp.linspace(0, 1, 100)
128
+ assert xp.get_backend() == "numpy"
129
+
130
+ # The previous global backend is active again here.
131
+ ```
132
+
133
+ The backend selection is process-wide shared state. Do not switch it
134
+ independently from multiple threads or async tasks; those changes can
135
+ interfere. A context manager is useful for sequential code, tests, and
136
+ notebooks.
137
+
138
+ ## Understand the two backend questions
139
+
140
+ The active backend controls which library CuNumpy exposes through its NumPy
141
+ like operations. The array backend reports where one particular array lives.
142
+ These can differ: changing the active backend does not convert arrays that
143
+ already exist.
144
+
145
+ ```python
146
+ xp.set_backend("numpy")
147
+ cpu_values = xp.arange(3)
148
+
149
+ gpu_values = xp.to_cupy(cpu_values) # explicit transfer
150
+ print(xp.get_backend()) # 'numpy'
151
+ print(xp.get_array_backend(gpu_values)) # 'cupy'
152
+ ```
153
+
154
+ Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when
155
+ dispatch should follow the array passed to a function. `get_array_module()`
156
+ returns the matching `array_api_compat` module, which is useful when writing
157
+ backend-generic functions:
158
+
159
+ ```python
160
+ def vector_norm(values):
161
+ array_xp = xp.get_array_module(values)
162
+ return array_xp.sqrt(array_xp.sum(values * values))
163
+ ```
164
+
165
+ ## Move data between CPU and GPU
166
+
167
+ Transfers are explicit so it is clear when data crosses the CPU/GPU boundary:
168
+
169
+ ```python
170
+ host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host)
171
+ device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device)
172
+ active = xp.to_cunumpy(host) # convert to the currently selected backend
173
+ ```
174
+
175
+ `to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises
176
+ `ImportError` when CuPy or a functional CUDA runtime is unavailable.
177
+ `to_cunumpy()` is useful at API boundaries where the consumer expects the
178
+ currently selected backend. It does not change the original array.
179
+
180
+ Avoid transferring data inside a tight loop. Keep intermediate arrays on one
181
+ backend and move only at boundaries such as file I/O, plotting, or a
182
+ CPU-only library call. For example:
183
+
184
+ ```python
185
+ with xp.use_backend("cupy"):
186
+ signal = xp.asarray(host_signal)
187
+ filtered = xp.fft.rfft(signal)
188
+ result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer
189
+ ```
190
+
191
+ To verify that a block, such as a time step, makes no transfer at all, count
192
+ them: `count_transfers()` records every `to_numpy()`, `to_cupy()` and
193
+ `to_cunumpy()` call that actually copies, mirror/staging refreshes, argument
194
+ conversion and kernel output copy-back, with call sites and payload byte counts.
195
+ Host kernel conversions and fallbacks have separate explanatory markers.
196
+ `assert_no_transfers()` rejects host/device movement and permits device-only
197
+ conversions. Only CuNumpy execution/conversion helpers are counted; forwarded
198
+ backend calls such as `xp.asarray()` and raw
199
+ `cupy.ndarray.get()` or `cupy.asarray()` calls need a profiler such as `nsys`.
200
+
201
+ ```python
202
+ with xp.profiling.count_transfers() as counter:
203
+ propagator(dt)
204
+
205
+ assert counter.to_host == counter.to_device == 0, counter.report()
206
+ print(counter.bytes_to_host, counter.bytes_to_device)
207
+ ```
208
+
209
+ ## Random numbers and dtypes
210
+
211
+ `get_rng(seed)` returns a random generator for the active backend. NumPy and
212
+ CuPy have similar generator APIs, though exact bit-for-bit sequences are not
213
+ guaranteed to match between libraries:
214
+
215
+ ```python
216
+ rng = xp.rng.get_rng(seed=42)
217
+ samples = rng.normal(size=1000)
218
+ ```
219
+
220
+ Use `default_float_dtype()` when code needs to explicitly request the active
221
+ backend's `float64` dtype rather than rely on Python scalar inference:
222
+
223
+ ```python
224
+ x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype())
225
+ ```
226
+
227
+ ## GPU selection and memory helpers
228
+
229
+ These helpers are useful for multi-GPU programs and for understanding CuPy's
230
+ memory behavior:
231
+
232
+ ```python
233
+ print("visible GPUs:", xp.cuda.device_count())
234
+ xp.cuda.set_device(0) # selects CUDA device 0 when CuPy is active
235
+ print("memory (free, total):", xp.cuda.memory_info())
236
+ ```
237
+
238
+ `set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA
239
+ hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA
240
+ cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the
241
+ active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a
242
+ round-robin convenience for MPI layouts where local ranks map contiguously to
243
+ GPUs. If your scheduler uses a different mapping, select the device directly.
244
+
245
+ For MPI programs with one rank per GPU, the startup sequence is:
246
+
247
+ 1. `bind_local_device()` selects the GPU from the node-local rank that the MPI
248
+ launcher exports (`local_rank()`) and creates its CUDA context. It runs
249
+ before MPI is initialized because a CUDA-aware MPI binds to the device that
250
+ is current at `MPI_Init`; without it, every rank of a node would use
251
+ device 0.
252
+ 2. `from mpi4py import MPI` initializes MPI.
253
+ 3. `require_cuda_aware_mpi()` (or `mpi_is_cuda_aware(comm)`) checks, with one
254
+ tiny device `Sendrecv` on every rank, that the MPI library can pass device
255
+ buffers at all. Passing CuPy arrays to a plain MPI build segfaults or
256
+ silently sends garbage; the check turns that into a clear error at
257
+ startup. It is a no-op on the NumPy backend.
258
+ 4. `synchronize_for_mpi(*buffers)` before every MPI call with device buffers:
259
+ kernels run asynchronously, and MPI would otherwise send a buffer a kernel
260
+ is still writing, without an error.
261
+
262
+ ```python
263
+ xp.set_backend("cupy")
264
+ xp.cuda.bind_local_device() # before MPI_Init
265
+ from mpi4py import MPI # MPI_Init
266
+
267
+ xp.mpi.require_cuda_aware_mpi() # once, on all ranks
268
+
269
+ xp.mpi.synchronize_for_mpi(send, recv)
270
+ MPI.COMM_WORLD.Sendrecv(send, dest, recvbuf=recv, source=source)
271
+ ```
272
+
273
+ CuPy caches released allocations in memory pools. This can make process-level
274
+ GPU memory appear occupied after arrays go out of scope. `free_memory()` asks
275
+ CuPy to release currently free cached blocks; it does not free memory still
276
+ referenced by live arrays.
277
+
278
+ `pin_memory(host_array)` makes a pinned host copy, which can improve transfer
279
+ throughput for workloads that explicitly manage asynchronous transfers.
280
+ `stream()` creates a non-blocking CuPy stream and yields it; it yields `None`
281
+ on NumPy. GPU work is asynchronous, so synchronize before reading results on
282
+ the host:
283
+
284
+ ```python
285
+ with xp.cuda.stream():
286
+ device = xp.to_cupy(host)
287
+ transformed = xp.fft.fft(device)
288
+
289
+ xp.synchronize()
290
+ result = xp.to_numpy(transformed)
291
+ ```
292
+
293
+ Because GPU work is asynchronous, a wall-clock timer around a kernel launch
294
+ measures the launch, not the kernel. `timed_region(name)` synchronizes the
295
+ device before reading the clock (on NumPy it is a plain timer), and
296
+ `nvtx_range(name)` marks a region so it shows up in `nsys`/Nsight; both are
297
+ no-ops or plain timers on NumPy, and `nvtx_range` also works as a decorator:
298
+
299
+ ```python
300
+ with xp.profiling.timed_region("fft") as timing:
301
+ transformed = xp.fft.fft(device)
302
+ print(timing.elapsed, timing.synced)
303
+
304
+
305
+ @xp.profiling.nvtx_range("step")
306
+ def step(dt): ...
307
+ ```
308
+
309
+ ## Use NumPy-only kernels with CuPy arrays
310
+
311
+ `PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is
312
+ needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function,
313
+ copies in-place output changes back to the device, and moves returned NumPy
314
+ arrays to CuPy. With NumPy inputs, the wrapper calls the function directly.
315
+ CuNumpy does not compile functions or import Pyccel for you.
316
+
317
+ ```python
318
+ import cunumpy as xp
319
+
320
+
321
+ def scale_in_place(values, factor):
322
+ values[:] *= factor
323
+ return values
324
+
325
+
326
+ scale = xp.kernels.PyccelKernel(scale_in_place, outputs=(0,))
327
+
328
+ with xp.use_backend("cupy"):
329
+ values = xp.arange(5, dtype=xp.float64)
330
+ returned = scale(values, 3.0)
331
+ xp.synchronize()
332
+ ```
333
+
334
+ By default every converted argument is copied back, because the wrapper cannot
335
+ know which arguments the kernel changed. `outputs=(0,)` declares that
336
+ positional argument 0 is written, avoiding unnecessary copy-back for
337
+ read-only inputs. For a keyword call, declare the keyword name, such as
338
+ `outputs=("out",)`. A wrong declaration can leave GPU output values stale.
339
+ The wrapper can also traverse arrays nested in lists, tuples, dictionaries,
340
+ and selected application objects; see the full [API reference](docs/source/api.md)
341
+ for `object_modules`, `is_array`, aliasing, and output declarations.
342
+
343
+ ## Write CUDA kernels next to host kernels
344
+
345
+ `CudaKernel` wraps a CUDA C kernel (compiled with NVRTC through
346
+ `cupy.RawKernel`) so that it is called with the same arguments as the host
347
+ kernel it mirrors. Thread counts default to the first array's leading shape
348
+ axes: one thread per row for 1D blocks, matching axes for 2D/3D blocks. Explicit
349
+ `n_threads`, `grid`, or a custom `n_threads_from` controls the launch when needed.
350
+ Arrays are never copied: they
351
+ must be C-contiguous CuPy arrays. The `extern "C" __global__` signature is
352
+ parsed once and every call is checked against it: Python scalars are cast to
353
+ the declared C types, and a wrong argument count, an array of the wrong dtype
354
+ or a non-contiguous view, or a scalar that does not fit its type raises instead
355
+ of silently producing wrong values.
356
+
357
+ `Kernel` pairs a host kernel with its CUDA kernel and calls the one matching
358
+ the active backend, so kernels can be ported to CUDA one at a time:
359
+
360
+ ```python
361
+ import cunumpy as xp
362
+
363
+ AXPY = r"""
364
+ extern "C" __global__
365
+ void axpy(double a, const double* x, double* y, int n) {
366
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
367
+ if (i < n) y[i] += a * x[i];
368
+ }
369
+ """
370
+
371
+
372
+ def axpy(a, x, y, n): # host version, e.g. compiled with Pyccel
373
+ for i in range(n):
374
+ y[i] += a * x[i]
375
+
376
+
377
+ kernel = xp.kernels.Kernel(axpy, xp.kernels.CudaKernel(AXPY, "axpy"))
378
+
379
+ with xp.use_backend("cupy"):
380
+ x = xp.arange(1000, dtype=xp.float64)
381
+ y = xp.zeros(1000)
382
+ kernel(2.0, x, y, 1000) # infer n_threads = x.shape[0], run the CUDA kernel
383
+ ```
384
+
385
+ On the CuPy backend, a `Kernel` without CUDA kernel raises
386
+ `NotImplementedError` (or, with `missing_cuda="fallback"`, runs the host kernel
387
+ through `PyccelKernel`, with host copies; `host_options` configure that
388
+ `PyccelKernel`). `KernelCatalog.from_package()` collects kernel pairs from a
389
+ package with one folder per kernel (`name/name_kernels.py` and
390
+ `name/name_cuda.cu`), and `catalog.compile_all()` compiles all CUDA kernels at
391
+ setup.
392
+
393
+ Groups of arguments can be passed as one: objects implementing
394
+ `__cuda_args__()` (see `CudaArguments`) are flattened into several kernel
395
+ arguments, and `CudaStruct` defines a C struct once (its C `declaration` and
396
+ the matching memory layout) and packs values into it, which the kernel takes
397
+ as one parameter:
398
+
399
+ ```python
400
+ Vec = xp.arguments.CudaStruct("Vec", [("data", "double*"), ("n", "int")])
401
+ scale = xp.kernels.CudaKernel(
402
+ Vec.declaration
403
+ + r"""
404
+ extern "C" __global__ void scale(Vec v, double a) {
405
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
406
+ if (i < v.n) v.data[i] *= a;
407
+ }""",
408
+ "scale",
409
+ structs=[Vec],
410
+ )
411
+ scale(Vec(data=y, n=y.size), 0.5, n_threads=y.size)
412
+ ```
413
+
414
+ When building such argument objects, `xp.as_device_array(value, dtype,
415
+ ndim=None)` applies the "reference or copy once" rule: a CuPy array that
416
+ already has the dtype and is C-contiguous is returned as it is, anything else
417
+ (a tuple such as `degree = (3, 3, 3)`, a host array, another dtype, a
418
+ non-contiguous view) is converted; dtype and layout changes can require separate
419
+ device copies. Call it once when the object is
420
+ built, not per kernel call; on the NumPy backend it raises, so host data is
421
+ never copied to the device implicitly.
422
+ When the host kernel takes such a group as one object too (e.g. a Pyccel class
423
+ holding NumPy arrays), write a CUDA class with the same constructor and
424
+ attributes (a `CudaStructArguments`, see below) and let the owner of the arrays
425
+ build the one for the active backend. CuNumpy passes argument objects through
426
+ as they are and never converts one form into the other:
427
+
428
+ ```python
429
+ args_class = CudaMarkerArguments if xp.is_gpu(markers) else MarkerArguments
430
+ particles.args_markers = args_class(markers, markers.shape[0])
431
+ kernel(particles.args_markers, dt) # host or CUDA kernel
432
+ ```
433
+
434
+ Kernels ported from pyccel index arrays like `markers[ip, j]`, which needs
435
+ shapes and strides rather than bare pointers. The shipped header
436
+ `cunumpy/array_view.cuh` (found by every `CudaKernel`) provides the strided
437
+ views `Array1D<T>` to `Array4D<T>`; a parameter or struct field of that type
438
+ takes a CuPy array, contiguous or not, and indexes `a(i, j)`. The struct can be
439
+ generated from the annotations of the pyccel argument class, so the Python
440
+ class is the one definition, and written to a header that a test keeps in sync:
441
+
442
+ ```python
443
+ class MarkerArguments:
444
+ def __init__(self, markers: "float[:, :]", n_markers: int, valid: "bool[:]"): ...
445
+
446
+
447
+ MarkerArgs = xp.arguments.CudaStruct.from_signature(MarkerArguments.__init__, "MarkerArgs")
448
+ MarkerArgs.to_header(
449
+ "marker_args.cuh"
450
+ ) # Array2D<double> markers; long long n_markers; ...
451
+ push = xp.kernels.CudaKernel(
452
+ r"""
453
+ #include "marker_args.cuh"
454
+ #include <cunumpy/index.cuh>
455
+ extern "C" __global__ void push(MarkerArgs m, double dt) {
456
+ CUNUMPY_THREAD_1D(ip, m.n_markers);
457
+ if (m.valid(ip)) m.markers(ip, 0) += dt * m.markers(ip, 3);
458
+ }""",
459
+ "push",
460
+ structs=[MarkerArgs],
461
+ include_dirs=["."],
462
+ )
463
+ push(
464
+ MarkerArgs(markers=markers, n_markers=markers.shape[0], valid=valid),
465
+ 0.1,
466
+ n_threads=markers.shape[0],
467
+ )
468
+ ```
469
+
470
+ Launches can be 1D to 3D (`n_threads=(nx, ny)`, `block_size=(16, 16)`) or use
471
+ an explicit `grid`, with dynamic shared memory (`shared_mem`) and a `stream`.
472
+ C++ function templates are instantiated with `template_args`, and
473
+ `CudaKernelVariants` caches kernels whose source is generated per variant
474
+ (e.g. per dimension and dtype). See the [API reference](docs/source/api.md) for
475
+ details.
476
+
477
+ Kernels run asynchronously, so a CUDA error (an illegal memory access, say)
478
+ normally surfaces at a later `.get()` or MPI call, far from the kernel that
479
+ caused it. In debug mode, enabled with `xp.cuda.set_cuda_debug(True)`, the
480
+ context manager `xp.cuda.cuda_debug()`, `CudaKernel(..., debug=True)` or the
481
+ environment variable `CUNUMPY_CUDA_DEBUG=1`, kernels are compiled with
482
+ `-lineinfo` and `-DCUNUMPY_BOUNDS_CHECK` and every launch is synchronized, so
483
+ the error is raised as a `RuntimeError` naming the kernel and its launch shape.
484
+ To find the faulting line and out-of-bounds accesses that do not crash, the
485
+ next step is NVIDIA's memory checker:
486
+ `CUNUMPY_CUDA_DEBUG=1 compute-sanitizer python -m pytest ...`.
487
+
488
+ ## Test kernel pairs
489
+
490
+ `cunumpy.kernel_testing` helps to test the ports with pytest. `assert_kernels_agree`
491
+ builds the arguments on both backends, runs the host and the CUDA kernel and
492
+ compares the arrays they wrote; with `catalog.parity_cases()`, one
493
+ parametrised test covers every ported kernel of a catalog. `BACKENDS` and
494
+ `requires_cupy` parametrize tests over the backends, skipping CuPy without a
495
+ GPU, and `device_function_kernel` wraps a `__device__` helper in an elementwise
496
+ kernel so it can be checked against its host version without writing a test
497
+ kernel:
498
+
499
+ ```python
500
+ import pytest
501
+ from cunumpy.kernel_testing import assert_kernels_agree
502
+
503
+
504
+ def make_args(backend, seed):
505
+ x = xp.to_cunumpy(np.random.default_rng(seed).random(1000))
506
+ return (x, 2.0, x.size)
507
+
508
+
509
+ @pytest.mark.parametrize("name, kernel", catalog.parity_cases())
510
+ def test_parity(name, kernel):
511
+ assert_kernels_agree(kernel, make_args, n_threads=1000)
512
+ ```
513
+
514
+ Accumulation kernels often write into a buffer that another library owns on
515
+ the host (a stencil vector's `_data`, exchanged over MPI). `DeviceMirror`
516
+ pairs that NumPy array with a device copy: `mirror.device` is the CuPy array
517
+ on the GPU and the host array itself on the CPU, `to_host()` copies back in
518
+ place (the host array keeps its identity) and `zero()` clears the buffer, so
519
+ the one transfer per accumulation is explicit. The shipped header
520
+ `cunumpy/atomic.cuh` (found automatically, see `cuda_include_dir()`) provides
521
+ `cunumpy_atomic_add()` and 2D/3D indexed variants for the many-threads-to-one-cell
522
+ writes:
523
+
524
+ ```python
525
+ mirror = xp.memory.DeviceMirror(vector._data)
526
+ mirror.zero()
527
+ accumulate(markers, mirror.device, n_threads=n_markers)
528
+ mirror.to_host() # vector._data holds the result on both backends
529
+ ```
530
+
531
+ ## Pyodide
532
+
533
+ CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA
534
+ there. Ordinary Python callables can be wrapped with `PyccelKernel` without
535
+ compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete
536
+ installation example and compatibility notes.
537
+
538
+ ## Documentation
539
+
540
+ The full documentation lives in [`docs/source`](docs/source/index.md) and is
541
+ published at <https://max-models.github.io/cunumpy/>:
542
+
543
+ * Getting started: [installation](docs/source/installation.md) and a
544
+ [quickstart](docs/source/quickstart.md) with a map of which guide covers what.
545
+ * User guide: [choosing a backend](docs/source/guides/backends.md),
546
+ [backend-agnostic code](docs/source/guides/portable-code.md),
547
+ [data movement](docs/source/guides/data-movement.md),
548
+ [devices, memory and streams](docs/source/guides/gpu-devices.md),
549
+ [MPI with one rank per GPU](docs/source/guides/mpi.md),
550
+ [timing and profiling](docs/source/guides/profiling.md).
551
+ * Porting kernels: [overview](docs/source/kernels/overview.md),
552
+ [`PyccelKernel`](docs/source/kernels/pyccel-kernel.md),
553
+ [`CudaKernel`](docs/source/kernels/cuda-kernel.md),
554
+ [`Kernel` and `KernelCatalog`](docs/source/kernels/dispatch.md),
555
+ [argument objects and structs](docs/source/kernels/arguments.md),
556
+ [accumulation kernels](docs/source/kernels/accumulation.md),
557
+ [debugging](docs/source/kernels/debugging.md),
558
+ [testing](docs/source/kernels/testing.md).
559
+ * [Worked examples](docs/source/examples/index.md),
560
+ [best practices](docs/source/best-practices.md),
561
+ [troubleshooting](docs/source/troubleshooting.md),
562
+ [Pyodide](docs/source/pyodide.md) and the
563
+ [API reference](docs/source/api.md).
564
+
565
+ ### For AI coding assistants
566
+
567
+ [`src/cunumpy/LLM_GUIDE.md`](src/cunumpy/LLM_GUIDE.md) is a compact,
568
+ self-contained guide to the API and its rules for LLM-based coding assistants.
569
+ It ships inside the installed package, so an assistant working in a project that
570
+ depends on CuNumpy can read it from `site-packages/cunumpy/LLM_GUIDE.md`, or
571
+ locate it with:
572
+
573
+ ```bash
574
+ python -c "import cunumpy, pathlib; print(pathlib.Path(cunumpy.__file__).parent / 'LLM_GUIDE.md')"
575
+ ```