fastmem 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,104 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here.
4
+ Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
5
+ versioning follows [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [0.1.0]
8
+
9
+ First public release.
10
+
11
+ ### Added
12
+
13
+ **Reading**
14
+ - `read(addr, size=None, as_=None, or_none=False)` - the single read entry
15
+ point. `size=None` uses the target pointer size, so `read(addr)` reads a
16
+ pointer directly. `as_` interprets the bytes as a number
17
+ (`'ptr'`, `'int'`, `'u32'`, `'f64'`, ...). `or_none=True` returns None
18
+ instead of raising.
19
+ - `read_many(addrs, size=None, as_=None, into=False, span=-1, threads=0)` -
20
+ batch reads. `into=True` returns one bytearray with no per-address
21
+ objects. `span` groups nearby addresses into windows read with a single
22
+ call. `threads` spreads work across threads.
23
+ - `read_region(base, size, chunks=0, or_none=False)` - contiguous read;
24
+ `chunks=N` yields a generator of blocks instead.
25
+ - `find(pattern, regions=None, align=1, limit=0, chunk_size=0)` - search
26
+ through `bytes.find` (C code), with alignment, result limit and
27
+ overlapping block reads.
28
+
29
+ **Virtual memory**
30
+ - `regions(min_size, max_size, readable_only, committed_only, start, end)` -
31
+ walk process memory through `VirtualQueryEx`.
32
+ - `query(addr)` - describe the region containing an address.
33
+ - `Region` - immutable region description: `base`, `size`, `state`,
34
+ `protect`, `type`, `end`, `committed`, `readable`, `is_image`,
35
+ `is_private`, `contains()`.
36
+
37
+ **C extension** (`src/fastmem/_fastmem.c`, optional)
38
+ - `batch_release` - batched read with the GIL released.
39
+ - `batch_pooled`, `batch_pooled_bytes` - the same read spread across a pool
40
+ of persistent worker threads. Workers are created on first use and park on
41
+ an event between calls; building them per call cost ~812 us on Windows,
42
+ more than the read itself. Measured 2.57x at 4 workers against 1.11x
43
+ before, and 1.80x instead of 0.20x on addresses too sparse to group.
44
+ - `pool_workers()`, `pool_shutdown()` - inspect and drain the pool. Shutdown
45
+ is registered with `Py_AtExit`, so no threads outlive the interpreter.
46
+ - `last_error()` - the Windows error code from the last failed read.
47
+ `ctypes.get_last_error()` cannot see inside the extension, so without this
48
+ every `ReadMemoryError` from a C-backed read reported code 0 and the
49
+ "partial copy" hint never appeared.
50
+ - `batch_bytes`, `batch_grouped_bytes` - return ready lists of bytes, with
51
+ no Python-side slicing.
52
+ - `batch_grouped` - grouped read writing into a caller buffer.
53
+ - `one`, `one_or`, `is_alive`.
54
+ - `backend.py` selects the implementation and falls back to pure Python
55
+ silently; `Process.backend()` and `backend.HAVE_C` report the active one.
56
+ - `setup.py` builds it when possible. Missing compiler or SDK is not fatal.
57
+
58
+ **Other**
59
+ - `read_many` accepts any iterable, including generators.
60
+ - Target bitness detection via `IsWow64Process2` → `IsWow64Process` →
61
+ system architecture; exposes `pointer_size`, `is_64bit`, `machine`,
62
+ `machine_name`, `page_size`.
63
+ - `is_alive()` through `GetExitCodeProcess`.
64
+ - Lazily formatted exception messages: building an error object costs
65
+ ~0.4 us instead of ~3 us for a formatted string.
66
+
67
+ ### Compatibility
68
+
69
+ - `as_='ptr'` follows the target bitness rather than always reading 8
70
+ bytes. Previously WOW64 targets hit unmapped memory.
71
+ - Minimum access rights by default:
72
+ `PROCESS_VM_READ | PROCESS_QUERY_LIMITED_INFORMATION` instead of
73
+ `PROCESS_QUERY_INFORMATION`, which works without elevation for processes
74
+ owned by the same user.
75
+ - Windows 7 through Server 2025; x86, x64, ARM64; Python 3.9-3.13; PyPy on
76
+ pure Python.
77
+ - Only public Windows APIs. No hardcoded structure offsets, no assembly.
78
+ Every optional API is feature-detected.
79
+ - Page size read from `GetSystemInfo`, never assumed.
80
+ - Protected processes (PPL) and PID 4 yield `ProcessOpenError` with
81
+ readable text instead of crashing.
82
+ - A process dying mid-read never raises from `or_none=True`, and plain
83
+ `read` raises `ProcessTerminatedError`.
84
+
85
+ ### Performance
86
+
87
+ Measured on a foreign process, microseconds per address, 8 bytes each:
88
+
89
+ | Operation | Python | C |
90
+ |---|---|---|
91
+ | batch into bytearray | 5.08 | 1.40 |
92
+ | batch into list[bytes] | 5.82 | 1.36 |
93
+ | batch into list[int] | 5.29 | 1.47 |
94
+ | grouping vs streaming | 5.12 | 0.076 |
95
+
96
+ The wins come from block size rather than language: one
97
+ `read_region(64 KiB)` costs ~22 us against ~150 ms for the same bytes read
98
+ address by address.
99
+
100
+ Threading is documented as 2.57x at 4 workers, not the 1.11x it measured
101
+ when the threads were built per call. Grouping still wins for dense
102
+ addresses; threads are for addresses too spread out to merge.
103
+
104
+ [0.1.0]: https://github.com/slitter-tech/fastmem/releases/tag/v0.1.0
fastmem-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 slitter-tech
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,20 @@
1
+ # Metadata
2
+ include README.md
3
+ include README_RU.md
4
+ include LICENSE
5
+ include CHANGELOG.md
6
+ include pyproject.toml
7
+ include setup.py
8
+
9
+ # The C source must ship in the sdist: without it there is nothing to
10
+ # compile when pip builds from source.
11
+ include src/fastmem/_fastmem.c
12
+
13
+ recursive-include tests *.py
14
+
15
+ # The prebuilt extension must not ship in the sdist.
16
+ exclude src/fastmem/*.pyd
17
+ exclude src/fastmem/*.so
18
+ prune build
19
+ prune dist
20
+ global-exclude __pycache__ *.py[cod]
fastmem-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,338 @@
1
+ Metadata-Version: 2.4
2
+ Name: fastmem
3
+ Version: 0.1.0
4
+ Summary: High-performance process memory reading for Windows (reverse engineering, debugging, forensics)
5
+ Author: wh1letr0e
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/slitter-tech/fastmem
8
+ Project-URL: Repository, https://github.com/slitter-tech/fastmem
9
+ Project-URL: Issues, https://github.com/slitter-tech/fastmem/issues
10
+ Project-URL: Changelog, https://github.com/slitter-tech/fastmem/blob/main/CHANGELOG.md
11
+ Keywords: windows,memory,read-process-memory,reverse-engineering,debugging,forensics,process,ctypes
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Environment :: Console
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Intended Audience :: Information Technology
16
+ Classifier: Operating System :: Microsoft :: Windows
17
+ Classifier: Operating System :: Microsoft :: Windows :: Windows 7
18
+ Classifier: Operating System :: Microsoft :: Windows :: Windows 10
19
+ Classifier: Operating System :: Microsoft :: Windows :: Windows 11
20
+ Classifier: Programming Language :: C
21
+ Classifier: Programming Language :: Python :: 3
22
+ Classifier: Programming Language :: Python :: 3.9
23
+ Classifier: Programming Language :: Python :: 3.10
24
+ Classifier: Programming Language :: Python :: 3.11
25
+ Classifier: Programming Language :: Python :: 3.12
26
+ Classifier: Programming Language :: Python :: 3.13
27
+ Classifier: Programming Language :: Python :: Implementation :: CPython
28
+ Classifier: Programming Language :: Python :: Implementation :: PyPy
29
+ Classifier: Topic :: Security
30
+ Classifier: Topic :: Software Development :: Debuggers
31
+ Classifier: Topic :: System :: Monitoring
32
+ Requires-Python: >=3.9
33
+ Description-Content-Type: text/markdown
34
+ License-File: LICENSE
35
+ Dynamic: license-file
36
+
37
+ [Русский](README_RU.md) | **English**
38
+
39
+ # fastmem
40
+
41
+ [![PyPI](https://img.shields.io/pypi/v/fastmem.svg)](https://pypi.org/project/fastmem/)
42
+ [![Python](https://img.shields.io/pypi/pyversions/fastmem.svg)](https://pypi.org/project/fastmem/)
43
+ [![CI](https://github.com/slitter-tech/fastmem/actions/workflows/ci.yml/badge.svg)](https://github.com/slitter-tech/fastmem/actions/workflows/ci.yml)
44
+ [![License](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
45
+
46
+ High-performance process memory reading for Windows. Pure `ctypes` and the
47
+ standard library. An optional C extension is built automatically when a
48
+ compiler is available, but is never required.
49
+
50
+ Built for reverse engineering, debugging, memory forensics and security
51
+ research.
52
+
53
+ ## Installation
54
+
55
+ ```bash
56
+ pip install fastmem
57
+ ```
58
+
59
+ From source, with a local C extension build:
60
+
61
+ ```bash
62
+ git clone https://github.com/slitter-tech/fastmem.git
63
+ cd fastmem
64
+ pip install -e .
65
+ ```
66
+
67
+ ## Quick start
68
+
69
+ ```python
70
+ from fastmem import Process
71
+
72
+ with Process(pid) as p:
73
+ print(p.machine_name, p.pointer_size * 8, "bit")
74
+ data = p.read(address, 32) # bytes
75
+ value = p.read(address, as_="float") # number
76
+ obj = p.read(address) # pointer_size bytes
77
+ ```
78
+
79
+ Hot loops use `or_none=True`, which returns `None` instead of raising:
80
+
81
+ ```python
82
+ hit = p.read(address, 8, or_none=True)
83
+ ```
84
+
85
+ ## API
86
+
87
+ Eight methods cover everything.
88
+
89
+ | Method | Result | Use for |
90
+ |---|---|---|
91
+ | `read(addr, size=None, as_=None, or_none=False)` | `bytes`, number or `None` | single reads |
92
+ | `read_many(addrs, size=None, as_=None, into=False, span=-1, threads=0)` | `list` or `bytearray` | **batches** |
93
+ | `read_region(base, size, chunks=0, or_none=False)` | `bytes` or generator | **scanners** |
94
+ | `regions(min_size=0, max_size=0, ...)` | generator of `Region` | virtual memory walk |
95
+ | `find(pattern, regions=None, align=1, limit=0, chunk_size=0)` | generator of addresses | **value search** |
96
+ | `query(addr)` | `Region` or `None` | describe one region |
97
+ | `is_alive()` | `bool` | liveness |
98
+ | `close()` | - | release the handle |
99
+
100
+ ### Reading
101
+
102
+ `size=None` means the target pointer size, so `read(addr)` reads a pointer
103
+ directly. `as_` interprets the bytes as a number:
104
+
105
+ | `as_` | Meaning |
106
+ |---|---|
107
+ | `None` | raw `bytes` |
108
+ | `'ptr'` | pointer, following the **target** bitness (8 on x64/ARM64, 4 on x86) |
109
+ | `'int'`, `'uint'` | unsigned 32-bit |
110
+ | `'i32'`, `'u32'` | signed / unsigned 32-bit |
111
+ | `'i64'`, `'u64'` | signed / unsigned 64-bit |
112
+ | `'i8'`, `'u8'`, `'i16'`, `'u16'` | narrow integers |
113
+ | `'float'`, `'f32'` | 32-bit float |
114
+ | `'double'`, `'f64'` | 64-bit float |
115
+ | `'long'`, `'ulong'` | unsigned 64-bit |
116
+
117
+ `int` and `long` are unsigned on purpose: memory reads are about raw values
118
+ (hashes, flags, pointers), and `0xDEADBEEF` coming back as `-559038737`
119
+ surprises everyone. Use `i32` / `i64` when you want signed interpretation.
120
+
121
+ ```python
122
+ with Process(pid) as p:
123
+ hp = p.read(addr, as_="float")
124
+ n = p.read(addr, as_="int")
125
+ ok = p.read(addr, as_="ptr", or_none=True)
126
+ ```
127
+
128
+ ### Batches
129
+
130
+ ```python
131
+ with Process(pid) as p:
132
+ values = p.read_many(addrs, as_="ptr") # list[int | None]
133
+ raw = p.read_many(addrs, into=True) # one bytearray, failures zeroed
134
+ ```
135
+
136
+ `span` groups addresses closer than `span` bytes and reads each group with
137
+ a single call. `-1` (default) uses four pages, `0` disables grouping.
138
+ Nearby addresses are the common case in real work - object fields, array
139
+ elements, list nodes - and grouping is up to **67x** faster than reading
140
+ them one at a time.
141
+
142
+ `threads` spreads the addresses over worker threads and requires the C
143
+ extension. The workers live in C, are created on first use and are reused
144
+ afterwards, so the cost is paid once rather than per call. Measured gain is
145
+ **1.5x at 2 threads and 2.6x at 4** on addresses too spread out for
146
+ grouping; grouping still wins whenever it applies, so pair `threads` with
147
+ `span=0` when the addresses are scattered.
148
+
149
+ ### Regions
150
+
151
+ ```python
152
+ with Process(pid) as p:
153
+ for reg in p.regions(max_size=64 << 20):
154
+ blob = p.read_region(reg.base, reg.size, or_none=True)
155
+ if blob:
156
+ scan(blob)
157
+
158
+ # stream a large region instead of holding it in memory
159
+ for chunk in p.read_region(base, size, chunks=1 << 20):
160
+ scan(chunk)
161
+ ```
162
+
163
+ ### Search
164
+
165
+ ```python
166
+ import struct
167
+ from fastmem import Process
168
+
169
+ needle = struct.pack("<Q", 0x1A2B3C4D5E6F7788)
170
+
171
+ with Process(pid) as p:
172
+ for addr in p.find(needle, align=8, limit=1000):
173
+ print(hex(addr))
174
+ ```
175
+
176
+ The search runs in C through `bytes.find`: locating a pattern inside a
177
+ megabyte region takes ~0.3 us, while a Python loop over the bytes is five
178
+ orders of magnitude slower. `chunk_size` reads in overlapping blocks so
179
+ hits straddling a boundary survive.
180
+
181
+ ## Target properties
182
+
183
+ Detected automatically via `IsWow64Process2` → `IsWow64Process` → system
184
+ architecture, so `as_='ptr'` follows the target, not the host.
185
+
186
+ ```python
187
+ p.pointer_size # 4 or 8
188
+ p.is_64bit # bool
189
+ p.machine_name # 'x86', 'x64', 'ARM64', ...
190
+ p.page_size # from GetSystemInfo
191
+ p.backend() # 'c-extension' or 'python-ctypes'
192
+ ```
193
+
194
+ ## Examples
195
+
196
+ Walk a pointer chain:
197
+
198
+ ```python
199
+ with Process(pid) as p:
200
+ node = p.read(root, as_="ptr")
201
+ chain = []
202
+ while node and len(chain) < 100:
203
+ chain.append(node)
204
+ node = p.read(node, or_none=True)
205
+ if node:
206
+ node = int.from_bytes(node, "little")
207
+ print("chain length:", len(chain))
208
+ ```
209
+
210
+ Dump every readable region:
211
+
212
+ ```python
213
+ with Process(pid) as p:
214
+ total = 0
215
+ for reg in p.regions(max_size=64 << 20):
216
+ blob = p.read_region(reg.base, reg.size, or_none=True)
217
+ if blob:
218
+ total += len(blob)
219
+ print("read", total, "bytes")
220
+ ```
221
+
222
+ ## Compatibility
223
+
224
+ * Windows 7, 8, 8.1, 10, 11, Server 2012+; x86, x64, ARM64.
225
+ * Python 3.9-3.13. PyPy works on pure Python.
226
+ * 32- and 64-bit targets, including WOW64.
227
+ * Protected processes (PPL) and PID 4 (System): insufficient rights give
228
+ `ProcessOpenError` with readable text and a Windows error code, never a
229
+ crash.
230
+ * A process dying mid-read: `or_none=True` yields `None`, plain `read`
231
+ raises `ProcessTerminatedError` (a `ReadMemoryError` subclass).
232
+ * Minimum access rights by default:
233
+ `PROCESS_VM_READ | PROCESS_QUERY_LIMITED_INFORMATION`, which works
234
+ without elevation for processes of the same user. Override with
235
+ `Process(pid, access=...)`.
236
+ * No hardcoded Windows structure offsets and no assembly: only public APIs,
237
+ each checked for availability.
238
+ * Page sizes are read from `GetSystemInfo`, never assumed.
239
+
240
+ ## Thread safety
241
+
242
+ A `Process` instance is **not** thread-safe (it holds a reusable buffer).
243
+ Open one `Process` per thread for manual parallel reads. `read_many` is the
244
+ exception: the worker pool is shared, guarded, and rebuilt on demand, so it
245
+ is safe to call from several Python threads at once.
246
+
247
+ ## Performance
248
+
249
+ Python 3.11, 4 cores, foreign process, 8 bytes per address. "Python" means
250
+ the same work through `ctypes` with the extension removed, so the columns
251
+ differ only in where the loop runs.
252
+
253
+ | Operation | Python | C | Speedup |
254
+ |---|---|---|---|
255
+ | `read()` in try/except, bad address | 10.46 | - | - |
256
+ | `read(or_none=True)`, bad address | 3.55 | - | - |
257
+ | `read()` success | 3.53 | - | - |
258
+ | `read(as_='int')` | 5.25 | - | - |
259
+ | batch into `bytearray` | 5.08 | 1.40 | 3.62x |
260
+ | batch into `list[bytes]` | 5.82 | 1.36 | 4.27x |
261
+ | batch into `list[int]` | 5.29 | 1.47 | 3.59x |
262
+ | **grouping vs streaming** | 5.12 | **0.076** | **67.2x** |
263
+
264
+ Large blocks, one API call each:
265
+
266
+ | | us | MB/s |
267
+ |---|---|---|
268
+ | `read_region(4 KiB)` | 9.7 | 421 |
269
+ | `read_region(64 KiB)` | 22.0 | 2977 |
270
+ | `read_region(1 MiB)` | 817 | 1284 |
271
+
272
+ ### Where the speed comes from
273
+
274
+ 1. **Block size matters more than language.** One `read_region(64 KiB)`
275
+ costs ~22 us; reading the same 64 KiB address by address costs ~150 ms.
276
+ 2. **Search belongs in C.** `bytes.find` beats a Python loop by five orders
277
+ of magnitude.
278
+ 3. **Allocation zeroes memory.** `(ctypes.c_char * n)()` memsets: ~400 us
279
+ for 1 MiB. A reusable buffer removes it.
280
+ 4. **Exceptions are expensive.** Formatting the error text cost 3.3 us; the
281
+ message is now built lazily in `__str__`.
282
+ 5. **C means one boundary crossing instead of a thousand.** A Python -> C
283
+ call costs ~0.6 us, and a 1000-address batch in Python pays that a
284
+ thousand times. The C loop crosses once.
285
+ 6. **Handing data to C has a price too.** `(ctypes.c_size_t * n)(*addrs)`
286
+ took 330 us for 2000 addresses; `array.array` fills the same buffer in
287
+ 66 us. Worth 12% of a threaded read on its own.
288
+
289
+ ### Threads
290
+
291
+ Measured on a foreign process, no grouping, microseconds per address:
292
+
293
+ | | 1 thread | 2 | 4 |
294
+ |---|---|---|---|
295
+ | C, worker pool | 1.28 | 0.85 (1.51x) | 0.50 (2.57x) |
296
+ | same, into a bytearray | 1.26 | 0.96 (1.32x) | 0.54 (2.32x) |
297
+ | same, sparse addresses | 1.24 | 0.93 (1.33x) | 0.69 (1.80x) |
298
+
299
+ Threads used to be a net loss here: building `threading.Thread` objects per
300
+ call cost ~812 us on Windows, more than the ~1700 us of work it was meant
301
+ to parallelise, so four threads measured 1.11x on dense input and 0.20x on
302
+ sparse input - worse than not threading at all. The workers now live in C,
303
+ are created on first use and park on an event between calls, which is what
304
+ turned those into 2.57x and 1.80x. Repeated runs put the four-thread figure
305
+ between 2.57x and 2.71x; absolute timings move with machine load, the
306
+ ratios do not.
307
+
308
+ A control group confirms the mechanism: a variant that holds the GIL gains
309
+ exactly nothing from four threads (1617 us against 1605 us for one), so the
310
+ limiter is the GIL, not contention inside the kernel. The pool releases it
311
+ for the whole read.
312
+
313
+ The calling thread takes a share of the work, so `threads=4` means five
314
+ readers. `threads` is capped at `len(addrs) - 1` for the same reason.
315
+
316
+ Prefer grouping for dense addresses: one read per window beats one read per
317
+ address by more than threads can recover. Reach for threads when the
318
+ addresses are too spread out for grouping to apply.
319
+
320
+ ## Building from source
321
+
322
+ ```bash
323
+ python setup.py build_ext --inplace # optional
324
+ python -m tests.test_fastmem # functional tests
325
+ python -m tests.benchmarks # benchmarks, own process
326
+ python -m tests.benchmarks --foreign # benchmarks, foreign process
327
+ ```
328
+
329
+ Building the extension needs MSVC and the Windows SDK. The SDK is not
330
+ always in the default location, so `setup.py` looks through `WindowsSdkDir`
331
+ and the usual paths including `G:\Windows Kits\10`. Some Python builds ship
332
+ no `pythonXY.lib` (built without `Py_ENABLE_SHARED`); in that case the
333
+ import library is generated from the DLL export table with `dumpbin` and
334
+ `lib.exe`.
335
+
336
+ ## License
337
+
338
+ MIT. See [LICENSE](LICENSE).