cunumpy 0.1.5__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
cunumpy-0.3.0/PKG-INFO ADDED
@@ -0,0 +1,356 @@
1
+ Metadata-Version: 2.4
2
+ Name: cunumpy
3
+ Version: 0.3.0
4
+ Summary: Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`.
5
+ Author: Max
6
+ Project-URL: Source, https://github.com/max-models/cunumpy
7
+ Keywords: python
8
+ Classifier: Development Status :: 3 - Alpha
9
+ Classifier: Programming Language :: Python :: 3 :: Only
10
+ Classifier: Programming Language :: Python :: 3.10
11
+ Classifier: Programming Language :: Python :: 3.11
12
+ Classifier: Programming Language :: Python :: 3.12
13
+ Classifier: Programming Language :: Python :: 3.13
14
+ Classifier: Programming Language :: Python :: 3.14
15
+ Requires-Python: >=3.10
16
+ Description-Content-Type: text/markdown
17
+ Requires-Dist: array-api-compat
18
+ Requires-Dist: numpy
19
+ Provides-Extra: dev
20
+ Requires-Dist: black[jupyter]; extra == "dev"
21
+ Requires-Dist: isort; extra == "dev"
22
+ Requires-Dist: cunumpy[docs,test-compiled]; extra == "dev"
23
+ Provides-Extra: docs
24
+ Requires-Dist: ipykernel; extra == "docs"
25
+ Requires-Dist: myst-parser; extra == "docs"
26
+ Requires-Dist: nbconvert; extra == "docs"
27
+ Requires-Dist: nbsphinx; extra == "docs"
28
+ Requires-Dist: jupyterlab; extra == "docs"
29
+ Requires-Dist: pre-commit; extra == "docs"
30
+ Requires-Dist: pyproject-fmt; extra == "docs"
31
+ Requires-Dist: sphinx; extra == "docs"
32
+ Requires-Dist: sphinx-book-theme; extra == "docs"
33
+ Provides-Extra: test
34
+ Requires-Dist: coverage; extra == "test"
35
+ Requires-Dist: pytest; extra == "test"
36
+ Provides-Extra: test-compiled
37
+ Requires-Dist: cunumpy[test]; extra == "test-compiled"
38
+ Requires-Dist: pyccel; extra == "test-compiled"
39
+
40
+ # CuNumpy
41
+
42
+ CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays
43
+ on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace
44
+ `import numpy as np` with `import cunumpy as xp`; the array operations you
45
+ already know then run on the selected backend.
46
+
47
+ ```python
48
+ import cunumpy as xp
49
+
50
+ values = xp.arange(5, dtype=xp.float64)
51
+ print(values * 2)
52
+ print(xp.get_backend()) # 'numpy' by default
53
+ ```
54
+
55
+ CuNumpy selects an array library for newly requested operations. It does not
56
+ move existing arrays just because the selected backend changes. This guide
57
+ covers backend selection, array movement, mixed CPU/GPU workflows, and the
58
+ helper APIs CuNumpy provides around NumPy and CuPy.
59
+
60
+ ## Install
61
+
62
+ ```bash
63
+ python -m pip install cunumpy
64
+ ```
65
+
66
+ NumPy and `array-api-compat` are installed as dependencies. To use a GPU,
67
+ install a CuPy package compatible with your CUDA environment as well. CuPy
68
+ installation depends on the CUDA version and platform; follow the CuPy
69
+ installation instructions for your system. CuNumpy does not install CUDA.
70
+
71
+ `array-api-compat` supplies NumPy and CuPy compatibility modules with more
72
+ consistent behavior for shared array operations. CuNumpy uses them internally;
73
+ your arrays remain ordinary NumPy or CuPy arrays. See [why CuNumpy uses
74
+ `array-api-compat`](docs/source/array-api-compat.md) for a plain-language
75
+ explanation and examples.
76
+
77
+ ## Choose a backend
78
+
79
+ CuNumpy starts with NumPy unless `ARRAY_BACKEND=cupy` is set before import.
80
+ You can also choose at runtime:
81
+
82
+ ```python
83
+ import cunumpy as xp
84
+
85
+ xp.set_backend("cupy")
86
+ print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional
87
+
88
+ values = xp.arange(5) # created by the active backend
89
+ ```
90
+
91
+ The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested
92
+ but unavailable or not functional, CuNumpy falls back to NumPy. Always check
93
+ `get_backend()` when the effective backend matters, such as when reporting
94
+ configuration or deciding whether GPU-specific work will happen.
95
+
96
+ Use `use_backend()` for a temporary selection. It restores the previous
97
+ selection when the block exits, including when an exception is raised:
98
+
99
+ ```python
100
+ with xp.use_backend("numpy"):
101
+ cpu_values = xp.linspace(0, 1, 100)
102
+ assert xp.get_backend() == "numpy"
103
+
104
+ # The previous global backend is active again here.
105
+ ```
106
+
107
+ The backend selection is process-wide shared state. Do not switch it
108
+ independently from multiple threads or async tasks; those changes can
109
+ interfere. A context manager is useful for sequential code, tests, and
110
+ notebooks.
111
+
112
+ ## Understand the two backend questions
113
+
114
+ The active backend controls which library CuNumpy exposes through its NumPy
115
+ like operations. The array backend reports where one particular array lives.
116
+ These can differ: changing the active backend does not convert arrays that
117
+ already exist.
118
+
119
+ ```python
120
+ xp.set_backend("numpy")
121
+ cpu_values = xp.arange(3)
122
+
123
+ gpu_values = xp.to_cupy(cpu_values) # explicit transfer
124
+ print(xp.get_backend()) # 'numpy'
125
+ print(xp.get_array_backend(gpu_values)) # 'cupy'
126
+ ```
127
+
128
+ Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when
129
+ dispatch should follow the array passed to a function. `get_array_module()`
130
+ returns the matching `array_api_compat` module, which is useful when writing
131
+ backend-generic functions:
132
+
133
+ ```python
134
+ def vector_norm(values):
135
+ array_xp = xp.get_array_module(values)
136
+ return array_xp.sqrt(array_xp.sum(values * values))
137
+ ```
138
+
139
+ ## Move data between CPU and GPU
140
+
141
+ Transfers are explicit so it is clear when data crosses the CPU/GPU boundary:
142
+
143
+ ```python
144
+ host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host)
145
+ device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device)
146
+ active = xp.to_cunumpy(host) # convert to the currently selected backend
147
+ ```
148
+
149
+ `to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises
150
+ `ImportError` when CuPy or a functional CUDA runtime is unavailable.
151
+ `to_cunumpy()` is useful at API boundaries where the consumer expects the
152
+ currently selected backend. It does not change the original array.
153
+
154
+ Avoid transferring data inside a tight loop. Keep intermediate arrays on one
155
+ backend and move only at boundaries such as file I/O, plotting, or a
156
+ CPU-only library call. For example:
157
+
158
+ ```python
159
+ with xp.use_backend("cupy"):
160
+ signal = xp.asarray(host_signal)
161
+ filtered = xp.fft.rfft(signal)
162
+ result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer
163
+ ```
164
+
165
+ ## Random numbers and dtypes
166
+
167
+ `get_rng(seed)` returns a random generator for the active backend. NumPy and
168
+ CuPy have similar generator APIs, though exact bit-for-bit sequences are not
169
+ guaranteed to match between libraries:
170
+
171
+ ```python
172
+ rng = xp.get_rng(seed=42)
173
+ samples = rng.normal(size=1000)
174
+ ```
175
+
176
+ Use `default_float_dtype()` when code needs to explicitly request the active
177
+ backend's `float64` dtype rather than rely on Python scalar inference:
178
+
179
+ ```python
180
+ x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype())
181
+ ```
182
+
183
+ ## GPU selection and memory helpers
184
+
185
+ These helpers are useful for multi-GPU programs and for understanding CuPy's
186
+ memory behavior:
187
+
188
+ ```python
189
+ print("visible GPUs:", xp.device_count())
190
+ xp.set_device(0) # selects CUDA device 0 when CuPy is active
191
+ print("memory (free, total):", xp.memory_info())
192
+ ```
193
+
194
+ `set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA
195
+ hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA
196
+ cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the
197
+ active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a
198
+ round-robin convenience for MPI layouts where local ranks map contiguously to
199
+ GPUs. If your scheduler uses a different mapping, select the device directly.
200
+
201
+ For MPI programs with one rank per GPU, `bind_local_device()` selects the GPU
202
+ from the node-local rank that the MPI launcher exports (`local_rank()`), so it
203
+ can run before MPI is initialized, as CUDA-aware MPI requires. Before passing
204
+ device buffers to MPI, call `synchronize_for_mpi(*buffers)`: kernels run
205
+ asynchronously, and MPI would otherwise send a buffer a kernel is still
206
+ writing, without an error.
207
+
208
+ ```python
209
+ xp.set_backend("cupy")
210
+ xp.bind_local_device() # before MPI_Init
211
+ from mpi4py import MPI
212
+
213
+ xp.synchronize_for_mpi(send, recv)
214
+ MPI.COMM_WORLD.Sendrecv(send, dest, recvbuf=recv, source=source)
215
+ ```
216
+
217
+ CuPy caches released allocations in memory pools. This can make process-level
218
+ GPU memory appear occupied after arrays go out of scope. `free_memory()` asks
219
+ CuPy to release currently free cached blocks; it does not free memory still
220
+ referenced by live arrays.
221
+
222
+ `pin_memory(host_array)` makes a pinned host copy, which can improve transfer
223
+ throughput for workloads that explicitly manage asynchronous transfers.
224
+ `stream()` creates a non-blocking CuPy stream and yields it; it yields `None`
225
+ on NumPy. GPU work is asynchronous, so synchronize before reading results on
226
+ the host:
227
+
228
+ ```python
229
+ with xp.stream():
230
+ device = xp.to_cupy(host)
231
+ transformed = xp.fft.fft(device)
232
+
233
+ xp.synchronize()
234
+ result = xp.to_numpy(transformed)
235
+ ```
236
+
237
+ ## Use NumPy-only kernels with CuPy arrays
238
+
239
+ `PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is
240
+ needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function,
241
+ copies in-place output changes back to the device, and moves returned NumPy
242
+ arrays to CuPy. With NumPy inputs, the wrapper calls the function directly.
243
+ CuNumpy does not compile functions or import Pyccel for you.
244
+
245
+ ```python
246
+ import cunumpy as xp
247
+
248
+
249
+ def scale_in_place(values, factor):
250
+ values[:] *= factor
251
+ return values
252
+
253
+
254
+ scale = xp.PyccelKernel(scale_in_place, outputs=(0,))
255
+
256
+ with xp.use_backend("cupy"):
257
+ values = xp.arange(5, dtype=xp.float64)
258
+ returned = scale(values, 3.0)
259
+ xp.synchronize()
260
+ ```
261
+
262
+ By default every converted argument is copied back, because the wrapper cannot
263
+ know which arguments the kernel changed. `outputs=(0,)` declares that
264
+ positional argument 0 is written, avoiding unnecessary copy-back for
265
+ read-only inputs. For a keyword call, declare the keyword name, such as
266
+ `outputs=("out",)`. A wrong declaration can leave GPU output values stale.
267
+ The wrapper can also traverse arrays nested in lists, tuples, dictionaries,
268
+ and selected application objects; see the full [API reference](docs/source/api.md)
269
+ for `object_modules`, `is_array`, aliasing, and output declarations.
270
+
271
+ ## Write CUDA kernels next to host kernels
272
+
273
+ `CudaKernel` wraps a CUDA C kernel (compiled with NVRTC through
274
+ `cupy.RawKernel`) so that it is called with the same arguments as the host
275
+ kernel it mirrors, plus the number of threads. Arrays are never copied: they
276
+ must be CuPy arrays. The `extern "C" __global__` signature is parsed once and
277
+ every call is checked against it: Python scalars are cast to the declared C
278
+ types, and a wrong argument count, an array of the wrong dtype, or a scalar
279
+ that does not fit its type raises instead of silently producing wrong values.
280
+
281
+ `Kernel` pairs a host kernel with its CUDA kernel and calls the one matching
282
+ the active backend, so kernels can be ported to CUDA one at a time:
283
+
284
+ ```python
285
+ import cunumpy as xp
286
+
287
+ AXPY = r"""
288
+ extern "C" __global__
289
+ void axpy(double a, const double* x, double* y, int n) {
290
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
291
+ if (i < n) y[i] += a * x[i];
292
+ }
293
+ """
294
+
295
+
296
+ def axpy(a, x, y, n): # host version, e.g. compiled with Pyccel
297
+ for i in range(n):
298
+ y[i] += a * x[i]
299
+
300
+
301
+ kernel = xp.Kernel(axpy, xp.CudaKernel(AXPY, "axpy"))
302
+
303
+ with xp.use_backend("cupy"):
304
+ x = xp.arange(1000, dtype=xp.float64)
305
+ y = xp.zeros(1000)
306
+ kernel(2.0, x, y, 1000, n_threads=1000) # runs the CUDA kernel
307
+ ```
308
+
309
+ On the CuPy backend, a `Kernel` without CUDA kernel raises
310
+ `NotImplementedError` (or, with `missing_cuda="fallback"`, runs the host kernel
311
+ through `PyccelKernel`, with host copies; `host_options` configure that
312
+ `PyccelKernel`). `KernelCatalog.from_package()` collects kernel pairs from a
313
+ package with one folder per kernel (`name/name_kernels.py` and
314
+ `name/name_cuda.cu`), and `catalog.compile_all()` compiles all CUDA kernels at
315
+ setup.
316
+
317
+ Groups of arguments can be passed as one: objects implementing
318
+ `__cuda_args__()` (see `CudaArguments`) are flattened into several kernel
319
+ arguments, and `CudaStruct` defines a C struct once (its C `declaration` and
320
+ the matching memory layout) and packs values into it, which the kernel takes
321
+ as one parameter:
322
+
323
+ ```python
324
+ Vec = xp.CudaStruct("Vec", [("data", "double*"), ("n", "int")])
325
+ scale = xp.CudaKernel(
326
+ Vec.declaration
327
+ + r"""
328
+ extern "C" __global__ void scale(Vec v, double a) {
329
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
330
+ if (i < v.n) v.data[i] *= a;
331
+ }""",
332
+ "scale",
333
+ structs=[Vec],
334
+ )
335
+ scale(Vec(data=y, n=y.size), 0.5, n_threads=y.size)
336
+ ```
337
+
338
+ Launches can be 1D to 3D (`n_threads=(nx, ny)`, `block_size=(16, 16)`) or use
339
+ an explicit `grid`, with dynamic shared memory (`shared_mem`) and a `stream`.
340
+ C++ function templates are instantiated with `template_args`, and
341
+ `CudaKernelVariants` caches kernels whose source is generated per variant
342
+ (e.g. per dimension and dtype). See the [API reference](docs/source/api.md) for
343
+ details.
344
+
345
+ ## Pyodide
346
+
347
+ CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA
348
+ there. Ordinary Python callables can be wrapped with `PyccelKernel` without
349
+ compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete
350
+ installation example and compatibility notes.
351
+
352
+ ## Documentation
353
+
354
+ The [user guide](docs/source/quickstart.md) explains common workflows. The
355
+ [API reference](docs/source/api.md) documents each helper and its behavior.
356
+ The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage.
@@ -0,0 +1,317 @@
1
+ # CuNumpy
2
+
3
+ CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays
4
+ on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace
5
+ `import numpy as np` with `import cunumpy as xp`; the array operations you
6
+ already know then run on the selected backend.
7
+
8
+ ```python
9
+ import cunumpy as xp
10
+
11
+ values = xp.arange(5, dtype=xp.float64)
12
+ print(values * 2)
13
+ print(xp.get_backend()) # 'numpy' by default
14
+ ```
15
+
16
+ CuNumpy selects an array library for newly requested operations. It does not
17
+ move existing arrays just because the selected backend changes. This guide
18
+ covers backend selection, array movement, mixed CPU/GPU workflows, and the
19
+ helper APIs CuNumpy provides around NumPy and CuPy.
20
+
21
+ ## Install
22
+
23
+ ```bash
24
+ python -m pip install cunumpy
25
+ ```
26
+
27
+ NumPy and `array-api-compat` are installed as dependencies. To use a GPU,
28
+ install a CuPy package compatible with your CUDA environment as well. CuPy
29
+ installation depends on the CUDA version and platform; follow the CuPy
30
+ installation instructions for your system. CuNumpy does not install CUDA.
31
+
32
+ `array-api-compat` supplies NumPy and CuPy compatibility modules with more
33
+ consistent behavior for shared array operations. CuNumpy uses them internally;
34
+ your arrays remain ordinary NumPy or CuPy arrays. See [why CuNumpy uses
35
+ `array-api-compat`](docs/source/array-api-compat.md) for a plain-language
36
+ explanation and examples.
37
+
38
+ ## Choose a backend
39
+
40
+ CuNumpy starts with NumPy unless `ARRAY_BACKEND=cupy` is set before import.
41
+ You can also choose at runtime:
42
+
43
+ ```python
44
+ import cunumpy as xp
45
+
46
+ xp.set_backend("cupy")
47
+ print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional
48
+
49
+ values = xp.arange(5) # created by the active backend
50
+ ```
51
+
52
+ The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested
53
+ but unavailable or not functional, CuNumpy falls back to NumPy. Always check
54
+ `get_backend()` when the effective backend matters, such as when reporting
55
+ configuration or deciding whether GPU-specific work will happen.
56
+
57
+ Use `use_backend()` for a temporary selection. It restores the previous
58
+ selection when the block exits, including when an exception is raised:
59
+
60
+ ```python
61
+ with xp.use_backend("numpy"):
62
+ cpu_values = xp.linspace(0, 1, 100)
63
+ assert xp.get_backend() == "numpy"
64
+
65
+ # The previous global backend is active again here.
66
+ ```
67
+
68
+ The backend selection is process-wide shared state. Do not switch it
69
+ independently from multiple threads or async tasks; those changes can
70
+ interfere. A context manager is useful for sequential code, tests, and
71
+ notebooks.
72
+
73
+ ## Understand the two backend questions
74
+
75
+ The active backend controls which library CuNumpy exposes through its NumPy
76
+ like operations. The array backend reports where one particular array lives.
77
+ These can differ: changing the active backend does not convert arrays that
78
+ already exist.
79
+
80
+ ```python
81
+ xp.set_backend("numpy")
82
+ cpu_values = xp.arange(3)
83
+
84
+ gpu_values = xp.to_cupy(cpu_values) # explicit transfer
85
+ print(xp.get_backend()) # 'numpy'
86
+ print(xp.get_array_backend(gpu_values)) # 'cupy'
87
+ ```
88
+
89
+ Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when
90
+ dispatch should follow the array passed to a function. `get_array_module()`
91
+ returns the matching `array_api_compat` module, which is useful when writing
92
+ backend-generic functions:
93
+
94
+ ```python
95
+ def vector_norm(values):
96
+ array_xp = xp.get_array_module(values)
97
+ return array_xp.sqrt(array_xp.sum(values * values))
98
+ ```
99
+
100
+ ## Move data between CPU and GPU
101
+
102
+ Transfers are explicit so it is clear when data crosses the CPU/GPU boundary:
103
+
104
+ ```python
105
+ host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host)
106
+ device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device)
107
+ active = xp.to_cunumpy(host) # convert to the currently selected backend
108
+ ```
109
+
110
+ `to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises
111
+ `ImportError` when CuPy or a functional CUDA runtime is unavailable.
112
+ `to_cunumpy()` is useful at API boundaries where the consumer expects the
113
+ currently selected backend. It does not change the original array.
114
+
115
+ Avoid transferring data inside a tight loop. Keep intermediate arrays on one
116
+ backend and move only at boundaries such as file I/O, plotting, or a
117
+ CPU-only library call. For example:
118
+
119
+ ```python
120
+ with xp.use_backend("cupy"):
121
+ signal = xp.asarray(host_signal)
122
+ filtered = xp.fft.rfft(signal)
123
+ result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer
124
+ ```
125
+
126
+ ## Random numbers and dtypes
127
+
128
+ `get_rng(seed)` returns a random generator for the active backend. NumPy and
129
+ CuPy have similar generator APIs, though exact bit-for-bit sequences are not
130
+ guaranteed to match between libraries:
131
+
132
+ ```python
133
+ rng = xp.get_rng(seed=42)
134
+ samples = rng.normal(size=1000)
135
+ ```
136
+
137
+ Use `default_float_dtype()` when code needs to explicitly request the active
138
+ backend's `float64` dtype rather than rely on Python scalar inference:
139
+
140
+ ```python
141
+ x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype())
142
+ ```
143
+
144
+ ## GPU selection and memory helpers
145
+
146
+ These helpers are useful for multi-GPU programs and for understanding CuPy's
147
+ memory behavior:
148
+
149
+ ```python
150
+ print("visible GPUs:", xp.device_count())
151
+ xp.set_device(0) # selects CUDA device 0 when CuPy is active
152
+ print("memory (free, total):", xp.memory_info())
153
+ ```
154
+
155
+ `set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA
156
+ hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA
157
+ cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the
158
+ active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a
159
+ round-robin convenience for MPI layouts where local ranks map contiguously to
160
+ GPUs. If your scheduler uses a different mapping, select the device directly.
161
+
162
+ For MPI programs with one rank per GPU, `bind_local_device()` selects the GPU
163
+ from the node-local rank that the MPI launcher exports (`local_rank()`), so it
164
+ can run before MPI is initialized, as CUDA-aware MPI requires. Before passing
165
+ device buffers to MPI, call `synchronize_for_mpi(*buffers)`: kernels run
166
+ asynchronously, and MPI would otherwise send a buffer a kernel is still
167
+ writing, without an error.
168
+
169
+ ```python
170
+ xp.set_backend("cupy")
171
+ xp.bind_local_device() # before MPI_Init
172
+ from mpi4py import MPI
173
+
174
+ xp.synchronize_for_mpi(send, recv)
175
+ MPI.COMM_WORLD.Sendrecv(send, dest, recvbuf=recv, source=source)
176
+ ```
177
+
178
+ CuPy caches released allocations in memory pools. This can make process-level
179
+ GPU memory appear occupied after arrays go out of scope. `free_memory()` asks
180
+ CuPy to release currently free cached blocks; it does not free memory still
181
+ referenced by live arrays.
182
+
183
+ `pin_memory(host_array)` makes a pinned host copy, which can improve transfer
184
+ throughput for workloads that explicitly manage asynchronous transfers.
185
+ `stream()` creates a non-blocking CuPy stream and yields it; it yields `None`
186
+ on NumPy. GPU work is asynchronous, so synchronize before reading results on
187
+ the host:
188
+
189
+ ```python
190
+ with xp.stream():
191
+ device = xp.to_cupy(host)
192
+ transformed = xp.fft.fft(device)
193
+
194
+ xp.synchronize()
195
+ result = xp.to_numpy(transformed)
196
+ ```
197
+
198
+ ## Use NumPy-only kernels with CuPy arrays
199
+
200
+ `PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is
201
+ needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function,
202
+ copies in-place output changes back to the device, and moves returned NumPy
203
+ arrays to CuPy. With NumPy inputs, the wrapper calls the function directly.
204
+ CuNumpy does not compile functions or import Pyccel for you.
205
+
206
+ ```python
207
+ import cunumpy as xp
208
+
209
+
210
+ def scale_in_place(values, factor):
211
+ values[:] *= factor
212
+ return values
213
+
214
+
215
+ scale = xp.PyccelKernel(scale_in_place, outputs=(0,))
216
+
217
+ with xp.use_backend("cupy"):
218
+ values = xp.arange(5, dtype=xp.float64)
219
+ returned = scale(values, 3.0)
220
+ xp.synchronize()
221
+ ```
222
+
223
+ By default every converted argument is copied back, because the wrapper cannot
224
+ know which arguments the kernel changed. `outputs=(0,)` declares that
225
+ positional argument 0 is written, avoiding unnecessary copy-back for
226
+ read-only inputs. For a keyword call, declare the keyword name, such as
227
+ `outputs=("out",)`. A wrong declaration can leave GPU output values stale.
228
+ The wrapper can also traverse arrays nested in lists, tuples, dictionaries,
229
+ and selected application objects; see the full [API reference](docs/source/api.md)
230
+ for `object_modules`, `is_array`, aliasing, and output declarations.
231
+
232
+ ## Write CUDA kernels next to host kernels
233
+
234
+ `CudaKernel` wraps a CUDA C kernel (compiled with NVRTC through
235
+ `cupy.RawKernel`) so that it is called with the same arguments as the host
236
+ kernel it mirrors, plus the number of threads. Arrays are never copied: they
237
+ must be CuPy arrays. The `extern "C" __global__` signature is parsed once and
238
+ every call is checked against it: Python scalars are cast to the declared C
239
+ types, and a wrong argument count, an array of the wrong dtype, or a scalar
240
+ that does not fit its type raises instead of silently producing wrong values.
241
+
242
+ `Kernel` pairs a host kernel with its CUDA kernel and calls the one matching
243
+ the active backend, so kernels can be ported to CUDA one at a time:
244
+
245
+ ```python
246
+ import cunumpy as xp
247
+
248
+ AXPY = r"""
249
+ extern "C" __global__
250
+ void axpy(double a, const double* x, double* y, int n) {
251
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
252
+ if (i < n) y[i] += a * x[i];
253
+ }
254
+ """
255
+
256
+
257
+ def axpy(a, x, y, n): # host version, e.g. compiled with Pyccel
258
+ for i in range(n):
259
+ y[i] += a * x[i]
260
+
261
+
262
+ kernel = xp.Kernel(axpy, xp.CudaKernel(AXPY, "axpy"))
263
+
264
+ with xp.use_backend("cupy"):
265
+ x = xp.arange(1000, dtype=xp.float64)
266
+ y = xp.zeros(1000)
267
+ kernel(2.0, x, y, 1000, n_threads=1000) # runs the CUDA kernel
268
+ ```
269
+
270
+ On the CuPy backend, a `Kernel` without CUDA kernel raises
271
+ `NotImplementedError` (or, with `missing_cuda="fallback"`, runs the host kernel
272
+ through `PyccelKernel`, with host copies; `host_options` configure that
273
+ `PyccelKernel`). `KernelCatalog.from_package()` collects kernel pairs from a
274
+ package with one folder per kernel (`name/name_kernels.py` and
275
+ `name/name_cuda.cu`), and `catalog.compile_all()` compiles all CUDA kernels at
276
+ setup.
277
+
278
+ Groups of arguments can be passed as one: objects implementing
279
+ `__cuda_args__()` (see `CudaArguments`) are flattened into several kernel
280
+ arguments, and `CudaStruct` defines a C struct once (its C `declaration` and
281
+ the matching memory layout) and packs values into it, which the kernel takes
282
+ as one parameter:
283
+
284
+ ```python
285
+ Vec = xp.CudaStruct("Vec", [("data", "double*"), ("n", "int")])
286
+ scale = xp.CudaKernel(
287
+ Vec.declaration
288
+ + r"""
289
+ extern "C" __global__ void scale(Vec v, double a) {
290
+ int i = blockDim.x * blockIdx.x + threadIdx.x;
291
+ if (i < v.n) v.data[i] *= a;
292
+ }""",
293
+ "scale",
294
+ structs=[Vec],
295
+ )
296
+ scale(Vec(data=y, n=y.size), 0.5, n_threads=y.size)
297
+ ```
298
+
299
+ Launches can be 1D to 3D (`n_threads=(nx, ny)`, `block_size=(16, 16)`) or use
300
+ an explicit `grid`, with dynamic shared memory (`shared_mem`) and a `stream`.
301
+ C++ function templates are instantiated with `template_args`, and
302
+ `CudaKernelVariants` caches kernels whose source is generated per variant
303
+ (e.g. per dimension and dtype). See the [API reference](docs/source/api.md) for
304
+ details.
305
+
306
+ ## Pyodide
307
+
308
+ CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA
309
+ there. Ordinary Python callables can be wrapped with `PyccelKernel` without
310
+ compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete
311
+ installation example and compatibility notes.
312
+
313
+ ## Documentation
314
+
315
+ The [user guide](docs/source/quickstart.md) explains common workflows. The
316
+ [API reference](docs/source/api.md) documents each helper and its behavior.
317
+ The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage.
@@ -5,22 +5,21 @@ requires = [ "setuptools", "wheel" ]
5
5
 
6
6
  [project]
7
7
  name = "cunumpy"
8
- version = "0.1.5"
8
+ version = "0.3.0"
9
9
  description = "Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`."
10
10
  readme = "README.md"
11
11
  keywords = [ "python" ]
12
12
  license = { file = "LICENSE.txt" }
13
13
  authors = [ { name = "Max" } ]
14
- requires-python = ">=3.8"
14
+ requires-python = ">=3.10"
15
15
  classifiers = [
16
16
  "Development Status :: 3 - Alpha",
17
17
  "Programming Language :: Python :: 3 :: Only",
18
- "Programming Language :: Python :: 3.8",
19
- "Programming Language :: Python :: 3.9",
20
18
  "Programming Language :: Python :: 3.10",
21
19
  "Programming Language :: Python :: 3.11",
22
20
  "Programming Language :: Python :: 3.12",
23
21
  "Programming Language :: Python :: 3.13",
22
+ "Programming Language :: Python :: 3.14",
24
23
  ]
25
24
  dependencies = [
26
25
  "array-api-compat",