fastmem 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- fastmem-0.1.0/CHANGELOG.md +104 -0
- fastmem-0.1.0/LICENSE +21 -0
- fastmem-0.1.0/MANIFEST.in +20 -0
- fastmem-0.1.0/PKG-INFO +338 -0
- fastmem-0.1.0/README.md +302 -0
- fastmem-0.1.0/README_RU.md +307 -0
- fastmem-0.1.0/pyproject.toml +74 -0
- fastmem-0.1.0/setup.cfg +4 -0
- fastmem-0.1.0/setup.py +313 -0
- fastmem-0.1.0/src/fastmem/__init__.py +42 -0
- fastmem-0.1.0/src/fastmem/_fastmem.c +1421 -0
- fastmem-0.1.0/src/fastmem/_winapi.py +304 -0
- fastmem-0.1.0/src/fastmem/backend.py +281 -0
- fastmem-0.1.0/src/fastmem/exceptions.py +82 -0
- fastmem-0.1.0/src/fastmem/process.py +957 -0
- fastmem-0.1.0/src/fastmem.egg-info/PKG-INFO +338 -0
- fastmem-0.1.0/src/fastmem.egg-info/SOURCES.txt +22 -0
- fastmem-0.1.0/src/fastmem.egg-info/dependency_links.txt +1 -0
- fastmem-0.1.0/src/fastmem.egg-info/not-zip-safe +1 -0
- fastmem-0.1.0/src/fastmem.egg-info/top_level.txt +1 -0
- fastmem-0.1.0/tests/__init__.py +31 -0
- fastmem-0.1.0/tests/benchmarks.py +338 -0
- fastmem-0.1.0/tests/test_fastmem.py +739 -0
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here.
|
|
4
|
+
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
5
|
+
versioning follows [Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [0.1.0]
|
|
8
|
+
|
|
9
|
+
First public release.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
**Reading**
|
|
14
|
+
- `read(addr, size=None, as_=None, or_none=False)` - the single read entry
|
|
15
|
+
point. `size=None` uses the target pointer size, so `read(addr)` reads a
|
|
16
|
+
pointer directly. `as_` interprets the bytes as a number
|
|
17
|
+
(`'ptr'`, `'int'`, `'u32'`, `'f64'`, ...). `or_none=True` returns None
|
|
18
|
+
instead of raising.
|
|
19
|
+
- `read_many(addrs, size=None, as_=None, into=False, span=-1, threads=0)` -
|
|
20
|
+
batch reads. `into=True` returns one bytearray with no per-address
|
|
21
|
+
objects. `span` groups nearby addresses into windows read with a single
|
|
22
|
+
call. `threads` spreads work across threads.
|
|
23
|
+
- `read_region(base, size, chunks=0, or_none=False)` - contiguous read;
|
|
24
|
+
`chunks=N` yields a generator of blocks instead.
|
|
25
|
+
- `find(pattern, regions=None, align=1, limit=0, chunk_size=0)` - search
|
|
26
|
+
through `bytes.find` (C code), with alignment, result limit and
|
|
27
|
+
overlapping block reads.
|
|
28
|
+
|
|
29
|
+
**Virtual memory**
|
|
30
|
+
- `regions(min_size, max_size, readable_only, committed_only, start, end)` -
|
|
31
|
+
walk process memory through `VirtualQueryEx`.
|
|
32
|
+
- `query(addr)` - describe the region containing an address.
|
|
33
|
+
- `Region` - immutable region description: `base`, `size`, `state`,
|
|
34
|
+
`protect`, `type`, `end`, `committed`, `readable`, `is_image`,
|
|
35
|
+
`is_private`, `contains()`.
|
|
36
|
+
|
|
37
|
+
**C extension** (`src/fastmem/_fastmem.c`, optional)
|
|
38
|
+
- `batch_release` - batched read with the GIL released.
|
|
39
|
+
- `batch_pooled`, `batch_pooled_bytes` - the same read spread across a pool
|
|
40
|
+
of persistent worker threads. Workers are created on first use and park on
|
|
41
|
+
an event between calls; building them per call cost ~812 us on Windows,
|
|
42
|
+
more than the read itself. Measured 2.57x at 4 workers against 1.11x
|
|
43
|
+
before, and 1.80x instead of 0.20x on addresses too sparse to group.
|
|
44
|
+
- `pool_workers()`, `pool_shutdown()` - inspect and drain the pool. Shutdown
|
|
45
|
+
is registered with `Py_AtExit`, so no threads outlive the interpreter.
|
|
46
|
+
- `last_error()` - the Windows error code from the last failed read.
|
|
47
|
+
`ctypes.get_last_error()` cannot see inside the extension, so without this
|
|
48
|
+
every `ReadMemoryError` from a C-backed read reported code 0 and the
|
|
49
|
+
"partial copy" hint never appeared.
|
|
50
|
+
- `batch_bytes`, `batch_grouped_bytes` - return ready lists of bytes, with
|
|
51
|
+
no Python-side slicing.
|
|
52
|
+
- `batch_grouped` - grouped read writing into a caller buffer.
|
|
53
|
+
- `one`, `one_or`, `is_alive`.
|
|
54
|
+
- `backend.py` selects the implementation and falls back to pure Python
|
|
55
|
+
silently; `Process.backend()` and `backend.HAVE_C` report the active one.
|
|
56
|
+
- `setup.py` builds it when possible. Missing compiler or SDK is not fatal.
|
|
57
|
+
|
|
58
|
+
**Other**
|
|
59
|
+
- `read_many` accepts any iterable, including generators.
|
|
60
|
+
- Target bitness detection via `IsWow64Process2` → `IsWow64Process` →
|
|
61
|
+
system architecture; exposes `pointer_size`, `is_64bit`, `machine`,
|
|
62
|
+
`machine_name`, `page_size`.
|
|
63
|
+
- `is_alive()` through `GetExitCodeProcess`.
|
|
64
|
+
- Lazily formatted exception messages: building an error object costs
|
|
65
|
+
~0.4 us instead of ~3 us for a formatted string.
|
|
66
|
+
|
|
67
|
+
### Compatibility
|
|
68
|
+
|
|
69
|
+
- `as_='ptr'` follows the target bitness rather than always reading 8
|
|
70
|
+
bytes. Previously WOW64 targets hit unmapped memory.
|
|
71
|
+
- Minimum access rights by default:
|
|
72
|
+
`PROCESS_VM_READ | PROCESS_QUERY_LIMITED_INFORMATION` instead of
|
|
73
|
+
`PROCESS_QUERY_INFORMATION`, which works without elevation for processes
|
|
74
|
+
owned by the same user.
|
|
75
|
+
- Windows 7 through Server 2025; x86, x64, ARM64; Python 3.9-3.13; PyPy on
|
|
76
|
+
pure Python.
|
|
77
|
+
- Only public Windows APIs. No hardcoded structure offsets, no assembly.
|
|
78
|
+
Every optional API is feature-detected.
|
|
79
|
+
- Page size read from `GetSystemInfo`, never assumed.
|
|
80
|
+
- Protected processes (PPL) and PID 4 yield `ProcessOpenError` with
|
|
81
|
+
readable text instead of crashing.
|
|
82
|
+
- A process dying mid-read never raises from `or_none=True`, and plain
|
|
83
|
+
`read` raises `ProcessTerminatedError`.
|
|
84
|
+
|
|
85
|
+
### Performance
|
|
86
|
+
|
|
87
|
+
Measured on a foreign process, microseconds per address, 8 bytes each:
|
|
88
|
+
|
|
89
|
+
| Operation | Python | C |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| batch into bytearray | 5.08 | 1.40 |
|
|
92
|
+
| batch into list[bytes] | 5.82 | 1.36 |
|
|
93
|
+
| batch into list[int] | 5.29 | 1.47 |
|
|
94
|
+
| grouping vs streaming | 5.12 | 0.076 |
|
|
95
|
+
|
|
96
|
+
The wins come from block size rather than language: one
|
|
97
|
+
`read_region(64 KiB)` costs ~22 us against ~150 ms for the same bytes read
|
|
98
|
+
address by address.
|
|
99
|
+
|
|
100
|
+
Threading is documented as 2.57x at 4 workers, not the 1.11x it measured
|
|
101
|
+
when the threads were built per call. Grouping still wins for dense
|
|
102
|
+
addresses; threads are for addresses too spread out to merge.
|
|
103
|
+
|
|
104
|
+
[0.1.0]: https://github.com/slitter-tech/fastmem/releases/tag/v0.1.0
|
fastmem-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 slitter-tech
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
# Metadata
|
|
2
|
+
include README.md
|
|
3
|
+
include README_RU.md
|
|
4
|
+
include LICENSE
|
|
5
|
+
include CHANGELOG.md
|
|
6
|
+
include pyproject.toml
|
|
7
|
+
include setup.py
|
|
8
|
+
|
|
9
|
+
# The C source must ship in the sdist: without it there is nothing to
|
|
10
|
+
# compile when pip builds from source.
|
|
11
|
+
include src/fastmem/_fastmem.c
|
|
12
|
+
|
|
13
|
+
recursive-include tests *.py
|
|
14
|
+
|
|
15
|
+
# The prebuilt extension must not ship in the sdist.
|
|
16
|
+
exclude src/fastmem/*.pyd
|
|
17
|
+
exclude src/fastmem/*.so
|
|
18
|
+
prune build
|
|
19
|
+
prune dist
|
|
20
|
+
global-exclude __pycache__ *.py[cod]
|
fastmem-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,338 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: fastmem
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: High-performance process memory reading for Windows (reverse engineering, debugging, forensics)
|
|
5
|
+
Author: wh1letr0e
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/slitter-tech/fastmem
|
|
8
|
+
Project-URL: Repository, https://github.com/slitter-tech/fastmem
|
|
9
|
+
Project-URL: Issues, https://github.com/slitter-tech/fastmem/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/slitter-tech/fastmem/blob/main/CHANGELOG.md
|
|
11
|
+
Keywords: windows,memory,read-process-memory,reverse-engineering,debugging,forensics,process,ctypes
|
|
12
|
+
Classifier: Development Status :: 4 - Beta
|
|
13
|
+
Classifier: Environment :: Console
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Intended Audience :: Information Technology
|
|
16
|
+
Classifier: Operating System :: Microsoft :: Windows
|
|
17
|
+
Classifier: Operating System :: Microsoft :: Windows :: Windows 7
|
|
18
|
+
Classifier: Operating System :: Microsoft :: Windows :: Windows 10
|
|
19
|
+
Classifier: Operating System :: Microsoft :: Windows :: Windows 11
|
|
20
|
+
Classifier: Programming Language :: C
|
|
21
|
+
Classifier: Programming Language :: Python :: 3
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
23
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
24
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
25
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
26
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
27
|
+
Classifier: Programming Language :: Python :: Implementation :: CPython
|
|
28
|
+
Classifier: Programming Language :: Python :: Implementation :: PyPy
|
|
29
|
+
Classifier: Topic :: Security
|
|
30
|
+
Classifier: Topic :: Software Development :: Debuggers
|
|
31
|
+
Classifier: Topic :: System :: Monitoring
|
|
32
|
+
Requires-Python: >=3.9
|
|
33
|
+
Description-Content-Type: text/markdown
|
|
34
|
+
License-File: LICENSE
|
|
35
|
+
Dynamic: license-file
|
|
36
|
+
|
|
37
|
+
[Русский](README_RU.md) | **English**
|
|
38
|
+
|
|
39
|
+
# fastmem
|
|
40
|
+
|
|
41
|
+
[](https://pypi.org/project/fastmem/)
|
|
42
|
+
[](https://pypi.org/project/fastmem/)
|
|
43
|
+
[](https://github.com/slitter-tech/fastmem/actions/workflows/ci.yml)
|
|
44
|
+
[](LICENSE)
|
|
45
|
+
|
|
46
|
+
High-performance process memory reading for Windows. Pure `ctypes` and the
|
|
47
|
+
standard library. An optional C extension is built automatically when a
|
|
48
|
+
compiler is available, but is never required.
|
|
49
|
+
|
|
50
|
+
Built for reverse engineering, debugging, memory forensics and security
|
|
51
|
+
research.
|
|
52
|
+
|
|
53
|
+
## Installation
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pip install fastmem
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
From source, with a local C extension build:
|
|
60
|
+
|
|
61
|
+
```bash
|
|
62
|
+
git clone https://github.com/slitter-tech/fastmem.git
|
|
63
|
+
cd fastmem
|
|
64
|
+
pip install -e .
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
## Quick start
|
|
68
|
+
|
|
69
|
+
```python
|
|
70
|
+
from fastmem import Process
|
|
71
|
+
|
|
72
|
+
with Process(pid) as p:
|
|
73
|
+
print(p.machine_name, p.pointer_size * 8, "bit")
|
|
74
|
+
data = p.read(address, 32) # bytes
|
|
75
|
+
value = p.read(address, as_="float") # number
|
|
76
|
+
obj = p.read(address) # pointer_size bytes
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
Hot loops use `or_none=True`, which returns `None` instead of raising:
|
|
80
|
+
|
|
81
|
+
```python
|
|
82
|
+
hit = p.read(address, 8, or_none=True)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
## API
|
|
86
|
+
|
|
87
|
+
Eight methods cover everything.
|
|
88
|
+
|
|
89
|
+
| Method | Result | Use for |
|
|
90
|
+
|---|---|---|
|
|
91
|
+
| `read(addr, size=None, as_=None, or_none=False)` | `bytes`, number or `None` | single reads |
|
|
92
|
+
| `read_many(addrs, size=None, as_=None, into=False, span=-1, threads=0)` | `list` or `bytearray` | **batches** |
|
|
93
|
+
| `read_region(base, size, chunks=0, or_none=False)` | `bytes` or generator | **scanners** |
|
|
94
|
+
| `regions(min_size=0, max_size=0, ...)` | generator of `Region` | virtual memory walk |
|
|
95
|
+
| `find(pattern, regions=None, align=1, limit=0, chunk_size=0)` | generator of addresses | **value search** |
|
|
96
|
+
| `query(addr)` | `Region` or `None` | describe one region |
|
|
97
|
+
| `is_alive()` | `bool` | liveness |
|
|
98
|
+
| `close()` | - | release the handle |
|
|
99
|
+
|
|
100
|
+
### Reading
|
|
101
|
+
|
|
102
|
+
`size=None` means the target pointer size, so `read(addr)` reads a pointer
|
|
103
|
+
directly. `as_` interprets the bytes as a number:
|
|
104
|
+
|
|
105
|
+
| `as_` | Meaning |
|
|
106
|
+
|---|---|
|
|
107
|
+
| `None` | raw `bytes` |
|
|
108
|
+
| `'ptr'` | pointer, following the **target** bitness (8 on x64/ARM64, 4 on x86) |
|
|
109
|
+
| `'int'`, `'uint'` | unsigned 32-bit |
|
|
110
|
+
| `'i32'`, `'u32'` | signed / unsigned 32-bit |
|
|
111
|
+
| `'i64'`, `'u64'` | signed / unsigned 64-bit |
|
|
112
|
+
| `'i8'`, `'u8'`, `'i16'`, `'u16'` | narrow integers |
|
|
113
|
+
| `'float'`, `'f32'` | 32-bit float |
|
|
114
|
+
| `'double'`, `'f64'` | 64-bit float |
|
|
115
|
+
| `'long'`, `'ulong'` | unsigned 64-bit |
|
|
116
|
+
|
|
117
|
+
`int` and `long` are unsigned on purpose: memory reads are about raw values
|
|
118
|
+
(hashes, flags, pointers), and `0xDEADBEEF` coming back as `-559038737`
|
|
119
|
+
surprises everyone. Use `i32` / `i64` when you want signed interpretation.
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
with Process(pid) as p:
|
|
123
|
+
hp = p.read(addr, as_="float")
|
|
124
|
+
n = p.read(addr, as_="int")
|
|
125
|
+
ok = p.read(addr, as_="ptr", or_none=True)
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
### Batches
|
|
129
|
+
|
|
130
|
+
```python
|
|
131
|
+
with Process(pid) as p:
|
|
132
|
+
values = p.read_many(addrs, as_="ptr") # list[int | None]
|
|
133
|
+
raw = p.read_many(addrs, into=True) # one bytearray, failures zeroed
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
`span` groups addresses closer than `span` bytes and reads each group with
|
|
137
|
+
a single call. `-1` (default) uses four pages, `0` disables grouping.
|
|
138
|
+
Nearby addresses are the common case in real work - object fields, array
|
|
139
|
+
elements, list nodes - and grouping is up to **67x** faster than reading
|
|
140
|
+
them one at a time.
|
|
141
|
+
|
|
142
|
+
`threads` spreads the addresses over worker threads and requires the C
|
|
143
|
+
extension. The workers live in C, are created on first use and are reused
|
|
144
|
+
afterwards, so the cost is paid once rather than per call. Measured gain is
|
|
145
|
+
**1.5x at 2 threads and 2.6x at 4** on addresses too spread out for
|
|
146
|
+
grouping; grouping still wins whenever it applies, so pair `threads` with
|
|
147
|
+
`span=0` when the addresses are scattered.
|
|
148
|
+
|
|
149
|
+
### Regions
|
|
150
|
+
|
|
151
|
+
```python
|
|
152
|
+
with Process(pid) as p:
|
|
153
|
+
for reg in p.regions(max_size=64 << 20):
|
|
154
|
+
blob = p.read_region(reg.base, reg.size, or_none=True)
|
|
155
|
+
if blob:
|
|
156
|
+
scan(blob)
|
|
157
|
+
|
|
158
|
+
# stream a large region instead of holding it in memory
|
|
159
|
+
for chunk in p.read_region(base, size, chunks=1 << 20):
|
|
160
|
+
scan(chunk)
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
### Search
|
|
164
|
+
|
|
165
|
+
```python
|
|
166
|
+
import struct
|
|
167
|
+
from fastmem import Process
|
|
168
|
+
|
|
169
|
+
needle = struct.pack("<Q", 0x1A2B3C4D5E6F7788)
|
|
170
|
+
|
|
171
|
+
with Process(pid) as p:
|
|
172
|
+
for addr in p.find(needle, align=8, limit=1000):
|
|
173
|
+
print(hex(addr))
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
The search runs in C through `bytes.find`: locating a pattern inside a
|
|
177
|
+
megabyte region takes ~0.3 us, while a Python loop over the bytes is five
|
|
178
|
+
orders of magnitude slower. `chunk_size` reads in overlapping blocks so
|
|
179
|
+
hits straddling a boundary survive.
|
|
180
|
+
|
|
181
|
+
## Target properties
|
|
182
|
+
|
|
183
|
+
Detected automatically via `IsWow64Process2` → `IsWow64Process` → system
|
|
184
|
+
architecture, so `as_='ptr'` follows the target, not the host.
|
|
185
|
+
|
|
186
|
+
```python
|
|
187
|
+
p.pointer_size # 4 or 8
|
|
188
|
+
p.is_64bit # bool
|
|
189
|
+
p.machine_name # 'x86', 'x64', 'ARM64', ...
|
|
190
|
+
p.page_size # from GetSystemInfo
|
|
191
|
+
p.backend() # 'c-extension' or 'python-ctypes'
|
|
192
|
+
```
|
|
193
|
+
|
|
194
|
+
## Examples
|
|
195
|
+
|
|
196
|
+
Walk a pointer chain:
|
|
197
|
+
|
|
198
|
+
```python
|
|
199
|
+
with Process(pid) as p:
|
|
200
|
+
node = p.read(root, as_="ptr")
|
|
201
|
+
chain = []
|
|
202
|
+
while node and len(chain) < 100:
|
|
203
|
+
chain.append(node)
|
|
204
|
+
node = p.read(node, or_none=True)
|
|
205
|
+
if node:
|
|
206
|
+
node = int.from_bytes(node, "little")
|
|
207
|
+
print("chain length:", len(chain))
|
|
208
|
+
```
|
|
209
|
+
|
|
210
|
+
Dump every readable region:
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
with Process(pid) as p:
|
|
214
|
+
total = 0
|
|
215
|
+
for reg in p.regions(max_size=64 << 20):
|
|
216
|
+
blob = p.read_region(reg.base, reg.size, or_none=True)
|
|
217
|
+
if blob:
|
|
218
|
+
total += len(blob)
|
|
219
|
+
print("read", total, "bytes")
|
|
220
|
+
```
|
|
221
|
+
|
|
222
|
+
## Compatibility
|
|
223
|
+
|
|
224
|
+
* Windows 7, 8, 8.1, 10, 11, Server 2012+; x86, x64, ARM64.
|
|
225
|
+
* Python 3.9-3.13. PyPy works on pure Python.
|
|
226
|
+
* 32- and 64-bit targets, including WOW64.
|
|
227
|
+
* Protected processes (PPL) and PID 4 (System): insufficient rights give
|
|
228
|
+
`ProcessOpenError` with readable text and a Windows error code, never a
|
|
229
|
+
crash.
|
|
230
|
+
* A process dying mid-read: `or_none=True` yields `None`, plain `read`
|
|
231
|
+
raises `ProcessTerminatedError` (a `ReadMemoryError` subclass).
|
|
232
|
+
* Minimum access rights by default:
|
|
233
|
+
`PROCESS_VM_READ | PROCESS_QUERY_LIMITED_INFORMATION`, which works
|
|
234
|
+
without elevation for processes of the same user. Override with
|
|
235
|
+
`Process(pid, access=...)`.
|
|
236
|
+
* No hardcoded Windows structure offsets and no assembly: only public APIs,
|
|
237
|
+
each checked for availability.
|
|
238
|
+
* Page sizes are read from `GetSystemInfo`, never assumed.
|
|
239
|
+
|
|
240
|
+
## Thread safety
|
|
241
|
+
|
|
242
|
+
A `Process` instance is **not** thread-safe (it holds a reusable buffer).
|
|
243
|
+
Open one `Process` per thread for manual parallel reads. `read_many` is the
|
|
244
|
+
exception: the worker pool is shared, guarded, and rebuilt on demand, so it
|
|
245
|
+
is safe to call from several Python threads at once.
|
|
246
|
+
|
|
247
|
+
## Performance
|
|
248
|
+
|
|
249
|
+
Python 3.11, 4 cores, foreign process, 8 bytes per address. "Python" means
|
|
250
|
+
the same work through `ctypes` with the extension removed, so the columns
|
|
251
|
+
differ only in where the loop runs.
|
|
252
|
+
|
|
253
|
+
| Operation | Python | C | Speedup |
|
|
254
|
+
|---|---|---|---|
|
|
255
|
+
| `read()` in try/except, bad address | 10.46 | - | - |
|
|
256
|
+
| `read(or_none=True)`, bad address | 3.55 | - | - |
|
|
257
|
+
| `read()` success | 3.53 | - | - |
|
|
258
|
+
| `read(as_='int')` | 5.25 | - | - |
|
|
259
|
+
| batch into `bytearray` | 5.08 | 1.40 | 3.62x |
|
|
260
|
+
| batch into `list[bytes]` | 5.82 | 1.36 | 4.27x |
|
|
261
|
+
| batch into `list[int]` | 5.29 | 1.47 | 3.59x |
|
|
262
|
+
| **grouping vs streaming** | 5.12 | **0.076** | **67.2x** |
|
|
263
|
+
|
|
264
|
+
Large blocks, one API call each:
|
|
265
|
+
|
|
266
|
+
| | us | MB/s |
|
|
267
|
+
|---|---|---|
|
|
268
|
+
| `read_region(4 KiB)` | 9.7 | 421 |
|
|
269
|
+
| `read_region(64 KiB)` | 22.0 | 2977 |
|
|
270
|
+
| `read_region(1 MiB)` | 817 | 1284 |
|
|
271
|
+
|
|
272
|
+
### Where the speed comes from
|
|
273
|
+
|
|
274
|
+
1. **Block size matters more than language.** One `read_region(64 KiB)`
|
|
275
|
+
costs ~22 us; reading the same 64 KiB address by address costs ~150 ms.
|
|
276
|
+
2. **Search belongs in C.** `bytes.find` beats a Python loop by five orders
|
|
277
|
+
of magnitude.
|
|
278
|
+
3. **Allocation zeroes memory.** `(ctypes.c_char * n)()` memsets: ~400 us
|
|
279
|
+
for 1 MiB. A reusable buffer removes it.
|
|
280
|
+
4. **Exceptions are expensive.** Formatting the error text cost 3.3 us; the
|
|
281
|
+
message is now built lazily in `__str__`.
|
|
282
|
+
5. **C means one boundary crossing instead of a thousand.** A Python -> C
|
|
283
|
+
call costs ~0.6 us, and a 1000-address batch in Python pays that a
|
|
284
|
+
thousand times. The C loop crosses once.
|
|
285
|
+
6. **Handing data to C has a price too.** `(ctypes.c_size_t * n)(*addrs)`
|
|
286
|
+
took 330 us for 2000 addresses; `array.array` fills the same buffer in
|
|
287
|
+
66 us. Worth 12% of a threaded read on its own.
|
|
288
|
+
|
|
289
|
+
### Threads
|
|
290
|
+
|
|
291
|
+
Measured on a foreign process, no grouping, microseconds per address:
|
|
292
|
+
|
|
293
|
+
| | 1 thread | 2 | 4 |
|
|
294
|
+
|---|---|---|---|
|
|
295
|
+
| C, worker pool | 1.28 | 0.85 (1.51x) | 0.50 (2.57x) |
|
|
296
|
+
| same, into a bytearray | 1.26 | 0.96 (1.32x) | 0.54 (2.32x) |
|
|
297
|
+
| same, sparse addresses | 1.24 | 0.93 (1.33x) | 0.69 (1.80x) |
|
|
298
|
+
|
|
299
|
+
Threads used to be a net loss here: building `threading.Thread` objects per
|
|
300
|
+
call cost ~812 us on Windows, more than the ~1700 us of work it was meant
|
|
301
|
+
to parallelise, so four threads measured 1.11x on dense input and 0.20x on
|
|
302
|
+
sparse input - worse than not threading at all. The workers now live in C,
|
|
303
|
+
are created on first use and park on an event between calls, which is what
|
|
304
|
+
turned those into 2.57x and 1.80x. Repeated runs put the four-thread figure
|
|
305
|
+
between 2.57x and 2.71x; absolute timings move with machine load, the
|
|
306
|
+
ratios do not.
|
|
307
|
+
|
|
308
|
+
A control group confirms the mechanism: a variant that holds the GIL gains
|
|
309
|
+
exactly nothing from four threads (1617 us against 1605 us for one), so the
|
|
310
|
+
limiter is the GIL, not contention inside the kernel. The pool releases it
|
|
311
|
+
for the whole read.
|
|
312
|
+
|
|
313
|
+
The calling thread takes a share of the work, so `threads=4` means five
|
|
314
|
+
readers. `threads` is capped at `len(addrs) - 1` for the same reason.
|
|
315
|
+
|
|
316
|
+
Prefer grouping for dense addresses: one read per window beats one read per
|
|
317
|
+
address by more than threads can recover. Reach for threads when the
|
|
318
|
+
addresses are too spread out for grouping to apply.
|
|
319
|
+
|
|
320
|
+
## Building from source
|
|
321
|
+
|
|
322
|
+
```bash
|
|
323
|
+
python setup.py build_ext --inplace # optional
|
|
324
|
+
python -m tests.test_fastmem # functional tests
|
|
325
|
+
python -m tests.benchmarks # benchmarks, own process
|
|
326
|
+
python -m tests.benchmarks --foreign # benchmarks, foreign process
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
Building the extension needs MSVC and the Windows SDK. The SDK is not
|
|
330
|
+
always in the default location, so `setup.py` looks through `WindowsSdkDir`
|
|
331
|
+
and the usual paths including `G:\Windows Kits\10`. Some Python builds ship
|
|
332
|
+
no `pythonXY.lib` (built without `Py_ENABLE_SHARED`); in that case the
|
|
333
|
+
import library is generated from the DLL export table with `dumpbin` and
|
|
334
|
+
`lib.exe`.
|
|
335
|
+
|
|
336
|
+
## License
|
|
337
|
+
|
|
338
|
+
MIT. See [LICENSE](LICENSE).
|