fastsimdjson 0.2.0__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- fastsimdjson-0.3.0/MANIFEST.in +2 -0
- fastsimdjson-0.3.0/PKG-INFO +446 -0
- fastsimdjson-0.3.0/README.md +415 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/pyproject.toml +2 -2
- fastsimdjson-0.3.0/setup.py +38 -0
- fastsimdjson-0.3.0/src/fastsimdjson.cpp +3602 -0
- fastsimdjson-0.3.0/src/fastsimdjson.egg-info/PKG-INFO +446 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/src/fastsimdjson.egg-info/SOURCES.txt +7 -1
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/src/fastsimdjson.egg-info/requires.txt +3 -0
- fastsimdjson-0.3.0/tests/test_dumps.py +417 -0
- fastsimdjson-0.3.0/tests/test_files.py +57 -0
- fastsimdjson-0.3.0/tests/test_stream.py +224 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/vendor/simdjson.cpp +244 -100
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/vendor/simdjson.h +4306 -1576
- fastsimdjson-0.3.0/vendor/zmij-LICENSE +21 -0
- fastsimdjson-0.3.0/vendor/zmij.cc +2263 -0
- fastsimdjson-0.3.0/vendor/zmij.h +868 -0
- fastsimdjson-0.2.0/MANIFEST.in +0 -2
- fastsimdjson-0.2.0/PKG-INFO +0 -252
- fastsimdjson-0.2.0/README.md +0 -224
- fastsimdjson-0.2.0/setup.py +0 -28
- fastsimdjson-0.2.0/src/fastsimdjson.cpp +0 -1390
- fastsimdjson-0.2.0/src/fastsimdjson.egg-info/PKG-INFO +0 -252
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/setup.cfg +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/src/fastsimdjson.egg-info/dependency_links.txt +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/src/fastsimdjson.egg-info/top_level.txt +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/tests/test_lazy.py +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/tests/test_loads.py +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/vendor/simdutf.cpp +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/vendor/simdutf.h +0 -0
- {fastsimdjson-0.2.0 → fastsimdjson-0.3.0}/vendor/simdutf_c.h +0 -0
|
@@ -0,0 +1,446 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: fastsimdjson
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: Fast JSON parsing for Python, built on simdjson
|
|
5
|
+
Author-email: Daniel Lemire <daniel@lemire.me>
|
|
6
|
+
License-Expression: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/simdjson/fastpysimdjson
|
|
8
|
+
Project-URL: Repository, https://github.com/simdjson/fastpysimdjson
|
|
9
|
+
Project-URL: Issues, https://github.com/simdjson/fastpysimdjson/issues
|
|
10
|
+
Keywords: json,simdjson,parser,simd,performance
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Intended Audience :: Developers
|
|
13
|
+
Classifier: Programming Language :: C++
|
|
14
|
+
Classifier: Programming Language :: Python :: 3
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
20
|
+
Classifier: Programming Language :: Python :: Implementation :: CPython
|
|
21
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
22
|
+
Requires-Python: >=3.10
|
|
23
|
+
Description-Content-Type: text/markdown
|
|
24
|
+
Provides-Extra: test
|
|
25
|
+
Requires-Dist: pytest; extra == "test"
|
|
26
|
+
Provides-Extra: bench
|
|
27
|
+
Requires-Dist: orjson; extra == "bench"
|
|
28
|
+
Requires-Dist: msgspec; extra == "bench"
|
|
29
|
+
Requires-Dist: pysimdjson; extra == "bench"
|
|
30
|
+
Requires-Dist: cysimdjson; extra == "bench"
|
|
31
|
+
|
|
32
|
+
# fastsimdjson
|
|
33
|
+
|
|
34
|
+
A Python binding for [simdjson](https://github.com/simdjson/simdjson) that
|
|
35
|
+
parses JSON into native Python objects (`dict`, `list`, `str`, `int`,
|
|
36
|
+
`float`, `bool`, `None`) and serializes them back. It is a drop-in
|
|
37
|
+
replacement for the `json` module's `loads`, `load`, `dumps` and `dump`, and
|
|
38
|
+
it can be over 3 times faster than the standard `json.loads` and
|
|
39
|
+
`json.dumps`. When you only need part of a document, its lazy `parse`
|
|
40
|
+
function is faster still. It also reads streams of documents (NDJSON, JSON
|
|
41
|
+
Lines). It writes JSON in two ways: `dumps` returns the same `str` as
|
|
42
|
+
`json.dumps`, and `dumpb` returns the same `bytes` as `orjson.dumps`, as fast
|
|
43
|
+
as orjson.
|
|
44
|
+
|
|
45
|
+
With pip:
|
|
46
|
+
|
|
47
|
+
```sh
|
|
48
|
+
pip install fastsimdjson
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
With uv:
|
|
52
|
+
|
|
53
|
+
```sh
|
|
54
|
+
uv pip install fastsimdjson # in a uv project: uv add fastsimdjson
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Wheels are available for Linux, macOS and Windows, for Python 3.10 to 3.14,
|
|
58
|
+
including free-threaded Python 3.14.
|
|
59
|
+
|
|
60
|
+
## Usage
|
|
61
|
+
|
|
62
|
+
### `loads`: the whole document
|
|
63
|
+
|
|
64
|
+
```python
|
|
65
|
+
import fastsimdjson
|
|
66
|
+
fastsimdjson.loads(b'{"a": [1, 2.5, "x", true, null]}')
|
|
67
|
+
# {'a': [1, 2.5, 'x', True, None]}
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
`loads(data)` accepts `bytes`, `bytearray`, `memoryview` and `str`, and
|
|
71
|
+
returns the same value as `json.loads`, with the same types and key order.
|
|
72
|
+
Integers that do not fit in 64 bits become exact Python ints. `NaN`,
|
|
73
|
+
`Infinity` and `-Infinity` are accepted, in any capitalization.
|
|
74
|
+
|
|
75
|
+
Invalid input raises `fastsimdjson.JSONDecodeError`, a subclass of
|
|
76
|
+
`json.JSONDecodeError`. When simdjson rejects a document, the same input is
|
|
77
|
+
parsed with `json.loads`. If that succeeds, `loads` returns its value (this
|
|
78
|
+
is how a number that overflows a double becomes `inf`). If it raises
|
|
79
|
+
`JSONDecodeError`, the exception is re-raised with Python's message and byte
|
|
80
|
+
position. Any other exception from `json.loads` propagates.
|
|
81
|
+
|
|
82
|
+
### `parse`: lazy views
|
|
83
|
+
|
|
84
|
+
When you need only part of a document, `parse` avoids building the rest.
|
|
85
|
+
It accepts the same inputs as `loads` and returns read-only views:
|
|
86
|
+
`fastsimdjson.Object` (a `Mapping`) and `fastsimdjson.Array` (a
|
|
87
|
+
`Sequence`). Values are converted when you access them; nested objects and
|
|
88
|
+
arrays are returned as views. A scalar root is returned as a plain value.
|
|
89
|
+
|
|
90
|
+
```python
|
|
91
|
+
doc = fastsimdjson.parse(open("twitter.json", "rb").read())
|
|
92
|
+
ids = [(s["id"], s["user"]["screen_name"]) for s in doc["statuses"]]
|
|
93
|
+
doc.at_pointer("/statuses/0/user/name") # JSON Pointer (RFC 6901)
|
|
94
|
+
doc["search_metadata"].as_dict() # convert a subtree, like loads
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
`Object` supports `obj[key]`, `get`, `in`, `len`, iteration over the keys,
|
|
98
|
+
`keys()`, `values()`, `items()` (iterators), `at_pointer` and `as_dict()`.
|
|
99
|
+
`Array` supports `arr[i]` (negative indexes and slices), `len`, iteration,
|
|
100
|
+
`at_pointer` and `as_list()`. Both work with `match` statements.
|
|
101
|
+
|
|
102
|
+
* A view keeps its document alive; the document owns its own buffers, so
|
|
103
|
+
it remains valid while other documents are parsed.
|
|
104
|
+
* A key lookup scans the object. With duplicate keys, lookups return the
|
|
105
|
+
first value, whereas `as_dict()` (like `json.loads`) keeps the last.
|
|
106
|
+
* Indexing an array walks it from the last index reached, so a loop over
|
|
107
|
+
`arr[i]` is linear; iteration is the fastest way to visit an array.
|
|
108
|
+
* Errors are handled as in `loads`. A document that simdjson rejects but
|
|
109
|
+
`json.loads` accepts (an overflowing number, an unpaired surrogate) is
|
|
110
|
+
returned as plain Python objects, as `loads` would return it.
|
|
111
|
+
|
|
112
|
+
### `dumps` and `dump`: like `json.dumps`
|
|
113
|
+
|
|
114
|
+
```python
|
|
115
|
+
fastsimdjson.dumps({"a": [1, 2.5, None]}) # '{"a": [1, 2.5, null]}'
|
|
116
|
+
fastsimdjson.dumps(obj, indent=2, sort_keys=True)
|
|
117
|
+
with open("out.json", "w", encoding="utf-8") as f:
|
|
118
|
+
fastsimdjson.dump(obj, f)
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
`dumps(obj, **kw)` takes the arguments of `json.dumps` and returns the same
|
|
122
|
+
`str`, character for character, including the float format (`repr`), the
|
|
123
|
+
escapes and the default separators. `ensure_ascii`, `indent`, `separators`,
|
|
124
|
+
`sort_keys`, `allow_nan` and `default` are handled in C. Everything else is
|
|
125
|
+
passed to `json.dumps` itself, which produces the result or raises its usual
|
|
126
|
+
exception: a `cls` argument or other encoder options, `skipkeys`, a circular
|
|
127
|
+
reference, `NaN` with `allow_nan=False`, a key or a value that `json` cannot
|
|
128
|
+
serialize. `dump(obj, fp, **kw)` writes `dumps(obj, **kw)` to `fp`.
|
|
129
|
+
|
|
130
|
+
### `dumpb`: like `orjson.dumps`
|
|
131
|
+
|
|
132
|
+
```python
|
|
133
|
+
fastsimdjson.dumpb({"a": [1, 2.5, None]}) # b'{"a":[1,2.5,null]}'
|
|
134
|
+
fastsimdjson.dumpb(obj, option=fastsimdjson.OPT_INDENT_2 | fastsimdjson.OPT_SORT_KEYS)
|
|
135
|
+
fastsimdjson.dumpb({1, 2}, default=sorted) # b'[1,2]'
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
`dumpb(obj, default=None, option=None)` has the arguments, the output and the
|
|
139
|
+
errors of `orjson.dumps` (orjson 3.12): `orjson.dumps(obj, ...)` can be
|
|
140
|
+
replaced by `fastsimdjson.dumpb(obj, ...)` for the types below. It returns
|
|
141
|
+
compact UTF-8 `bytes`, writes floats as orjson does (`1e-6`, `1e+16`, `NaN`
|
|
142
|
+
and infinities as `null`), and raises `TypeError` with orjson's messages:
|
|
143
|
+
integers beyond 64 bits, a dict key that is not a `str`, invalid UTF-8 (a lone
|
|
144
|
+
surrogate), nesting deeper than 254, a type it cannot serialize. `orjson.JSONEncodeError` is a subclass of `TypeError`, so
|
|
145
|
+
`except TypeError` catches the errors of both.
|
|
146
|
+
|
|
147
|
+
It serializes `str`, `int`, `float`, `bool`, `None`, `dict`, `list`, `tuple`,
|
|
148
|
+
enums, and subclasses of `str`, `int`, `dict` and `list`. Anything else goes
|
|
149
|
+
to `default`, as in orjson: its result is serialized in place of the object,
|
|
150
|
+
and an exception it raises becomes the `__cause__` of the `TypeError`. The
|
|
151
|
+
options are exported under orjson's names and values (`OPT_APPEND_NEWLINE`,
|
|
152
|
+
`OPT_INDENT_2`, `OPT_NON_STR_KEYS`, `OPT_PASSTHROUGH_SUBCLASS`,
|
|
153
|
+
`OPT_SORT_KEYS`, `OPT_STRICT_INTEGER`, ...). Unlike orjson, `dumpb` does not
|
|
154
|
+
serialize dataclasses, `datetime`, `date`, `time`, `UUID`, numpy arrays or
|
|
155
|
+
`orjson.Fragment` itself: they go to `default`, as if
|
|
156
|
+
`OPT_PASSTHROUGH_DATACLASS` and `OPT_PASSTHROUGH_DATETIME` were set, and the
|
|
157
|
+
options that concern only these types are accepted and have no effect.
|
|
158
|
+
|
|
159
|
+
### Files
|
|
160
|
+
|
|
161
|
+
```python
|
|
162
|
+
with open("data.json", "rb") as f:
|
|
163
|
+
doc = fastsimdjson.load(f) # like json.load: f.read(), then loads
|
|
164
|
+
doc = fastsimdjson.load_file("data.json")
|
|
165
|
+
view = fastsimdjson.parse_file("data.json") # lazy, like parse
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
`load_file(path)` and `parse_file(path)` accept a `str`, `bytes` or
|
|
169
|
+
`os.PathLike` path and raise `OSError` (e.g. `FileNotFoundError`) when the
|
|
170
|
+
file cannot be read.
|
|
171
|
+
|
|
172
|
+
### `loads_many` and `parse_many`: streams of documents
|
|
173
|
+
|
|
174
|
+
```python
|
|
175
|
+
for record in fastsimdjson.loads_many(open("log.ndjson", "rb").read()):
|
|
176
|
+
...
|
|
177
|
+
for view in fastsimdjson.parse_many(data): # lazy views, like parse
|
|
178
|
+
...
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
Both return an iterator over the documents of `data` (`bytes`, `bytearray`,
|
|
182
|
+
`memoryview` or `str`). The `format` keyword selects how documents are
|
|
183
|
+
separated:
|
|
184
|
+
|
|
185
|
+
| `format` | input |
|
|
186
|
+
|---|---|
|
|
187
|
+
| `"whitespace"` (default) | documents separated by white space, including NDJSON and JSON Lines |
|
|
188
|
+
| `"lines"` | one document per line (NDJSON, JSON Lines) |
|
|
189
|
+
| `"json_seq"` | RFC 7464 JSON text sequences (each document preceded by `\x1e`) |
|
|
190
|
+
| `"comma"` | documents separated by commas: `{...}, {...}` (simdjson also accepts white space between them) |
|
|
191
|
+
| `"array"` | the elements of one array: `[{...}, {...}]` |
|
|
192
|
+
|
|
193
|
+
simdjson parses the input in batches (`batch_size`, 1 MB by default); a
|
|
194
|
+
larger document is handled automatically. In every format, documents that
|
|
195
|
+
simdjson rejects are handled as in `loads`: a document that `json` accepts is
|
|
196
|
+
returned, otherwise `JSONDecodeError` reports `json`'s message and the
|
|
197
|
+
position in the whole input. A truncated last document is an error. The
|
|
198
|
+
views returned by `parse_many` remain valid after the iterator moves on.
|
|
199
|
+
|
|
200
|
+
### `release`
|
|
201
|
+
|
|
202
|
+
`release()` frees the simdjson parser and the string caches kept by the
|
|
203
|
+
calling thread. Views returned by `parse` remain valid.
|
|
204
|
+
|
|
205
|
+
## Build and test
|
|
206
|
+
|
|
207
|
+
Python 3.10 or newer, and a C++17 compiler (clang, GCC or MSVC). The simdjson
|
|
208
|
+
5.0.2 and simdutf 9.2.1 amalgamations, and zmij 1.2 (shortest float
|
|
209
|
+
formatting, MIT license), are already in `vendor/`.
|
|
210
|
+
|
|
211
|
+
pip:
|
|
212
|
+
|
|
213
|
+
```sh
|
|
214
|
+
python3 -m venv .venv
|
|
215
|
+
. .venv/bin/activate
|
|
216
|
+
python -m pip install -U pip setuptools
|
|
217
|
+
python -m pip install -e ".[test]"
|
|
218
|
+
pytest tests
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
uv:
|
|
222
|
+
|
|
223
|
+
```sh
|
|
224
|
+
uv venv
|
|
225
|
+
. .venv/bin/activate
|
|
226
|
+
uv pip install -e ".[test]"
|
|
227
|
+
pytest tests
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
Either one builds the extension and makes `import fastsimdjson` work in the
|
|
231
|
+
virtualenv. uv uses the `setuptools` build requirement from `pyproject.toml`,
|
|
232
|
+
so it does not need a separate setuptools install for this path.
|
|
233
|
+
|
|
234
|
+
To compile the extension in the tree instead, install setuptools and pytest
|
|
235
|
+
into the same virtualenv, then:
|
|
236
|
+
|
|
237
|
+
```sh
|
|
238
|
+
python -m pip install setuptools pytest
|
|
239
|
+
python setup.py build_ext --inplace
|
|
240
|
+
PYTHONPATH=src python -m pytest tests
|
|
241
|
+
```
|
|
242
|
+
|
|
243
|
+
```sh
|
|
244
|
+
uv pip install setuptools pytest
|
|
245
|
+
python setup.py build_ext --inplace
|
|
246
|
+
PYTHONPATH=src python -m pytest tests
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
Recent setuptools copies the `.so` next to `src/fastsimdjson.cpp`, which is
|
|
250
|
+
why `PYTHONPATH=src` is required for the in-place build.
|
|
251
|
+
|
|
252
|
+
`tests/test_loads.py` compares `loads` with `json.loads` on types and key
|
|
253
|
+
order; `tests/test_lazy.py` checks the views returned by `parse` the same way,
|
|
254
|
+
`tests/test_dumps.py` compares `dumps` with `json.dumps` and `dumpb` with
|
|
255
|
+
orjson (output and exceptions; the orjson comparisons are skipped when orjson
|
|
256
|
+
is not installed, and recorded orjson results are checked either way), and
|
|
257
|
+
`tests/test_stream.py` and `tests/test_files.py` cover streams and files. It
|
|
258
|
+
covers scalars, integers past 64 bits, UTF-8 strings at every length from 0 to
|
|
259
|
+
199, the key cache, random documents, rejected input, deep
|
|
260
|
+
nesting, padding at a page boundary, a saturated array count, reference
|
|
261
|
+
counts, and release of a parser that has grown past 64 MB. The corpus test
|
|
262
|
+
is skipped until simdjson-data is checked out beside the project:
|
|
263
|
+
|
|
264
|
+
```sh
|
|
265
|
+
git clone --depth 1 https://github.com/simdjson/simdjson-data.git
|
|
266
|
+
pytest tests
|
|
267
|
+
# or: JSONDIR=/path/to/jsonexamples pytest tests
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
The suite builds an ~80 MB document and a list of 16,777,221 integers, so
|
|
271
|
+
give it some RAM.
|
|
272
|
+
|
|
273
|
+
The benchmark scripts need simdjson-data and the `bench` extra (orjson,
|
|
274
|
+
msgspec, pysimdjson, cysimdjson). `bench.py` times `loads` against
|
|
275
|
+
`json.loads` and orjson, `bench_lazy.py` times `parse` against pysimdjson and
|
|
276
|
+
cysimdjson, `bench_dumps.py` times `dumps` against `json.dumps` and `dumpb`
|
|
277
|
+
against orjson and msgspec, and `bench_many.py` times `loads_many`.
|
|
278
|
+
|
|
279
|
+
pip:
|
|
280
|
+
|
|
281
|
+
```sh
|
|
282
|
+
python -m pip install -e ".[bench]"
|
|
283
|
+
python bench.py
|
|
284
|
+
python bench_lazy.py
|
|
285
|
+
python bench_dumps.py
|
|
286
|
+
python bench_many.py
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
uv:
|
|
290
|
+
|
|
291
|
+
```sh
|
|
292
|
+
uv pip install -e ".[bench]"
|
|
293
|
+
python bench.py
|
|
294
|
+
python bench_lazy.py
|
|
295
|
+
python bench_dumps.py
|
|
296
|
+
python bench_many.py
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
## Benchmarks
|
|
300
|
+
|
|
301
|
+
Intel Xeon Gold 6548N (Emerald Rapids), one core, Python 3.14.6,
|
|
302
|
+
fastsimdjson 0.3.0 (simdjson 5.0.2), the 22 files of
|
|
303
|
+
[simdjson-data](https://github.com/simdjson/simdjson-data). The scripts are in
|
|
304
|
+
this repository and in
|
|
305
|
+
[the blog repository](https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/pysimdjson).
|
|
306
|
+
|
|
307
|
+
### Whole documents
|
|
308
|
+
|
|
309
|
+
Each parser produces the whole document as Python objects. Speed is the
|
|
310
|
+
geometric mean over the 22 files (higher is better).
|
|
311
|
+
|
|
312
|
+
<picture>
|
|
313
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_loads_dark.png">
|
|
314
|
+
<img src="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_loads.png" width="80%" alt="Parsing JSON in Python: fastsimdjson loads 0.78 GB/s, orjson 0.60, msgspec 0.53, pysimdjson 0.44, cysimdjson 0.43, ujson 0.37, python-rapidjson 0.24, simplejson 0.23, json 0.22">
|
|
315
|
+
</picture>
|
|
316
|
+
|
|
317
|
+
| parser | GB/s | vs `json.loads` |
|
|
318
|
+
|---|---:|---:|
|
|
319
|
+
| json (standard library) | 0.22 | 1.00× |
|
|
320
|
+
| simplejson 4.2.0 | 0.23 | 1.05× |
|
|
321
|
+
| python-rapidjson 1.25 | 0.24 | 1.10× |
|
|
322
|
+
| ujson 6.0.0 | 0.37 | 1.66× |
|
|
323
|
+
| cysimdjson 26.27 | 0.43 | 1.96× |
|
|
324
|
+
| pysimdjson 7.0.2 | 0.44 | 1.98× |
|
|
325
|
+
| msgspec 0.22.0 | 0.53 | 2.42× |
|
|
326
|
+
| orjson 3.12.0 | 0.60 | 2.74× |
|
|
327
|
+
| **fastsimdjson `loads`** | **0.78** | **3.53×** |
|
|
328
|
+
|
|
329
|
+
fastsimdjson is the fastest on 21 of the 22 files, and ties with orjson on
|
|
330
|
+
`numbers.json`, an array of floating-point numbers (within 2%). On numbers,
|
|
331
|
+
both spend most of their time creating Python floats, which costs the same in
|
|
332
|
+
both: simdjson parses the numbers of these files 15% to 30% faster than
|
|
333
|
+
orjson's parser (yyjson), but that is a small part of the total. Part of the
|
|
334
|
+
gain comes from pausing the garbage collector while the objects are built.
|
|
335
|
+
Timing each call on its own, fastsimdjson's lead over orjson is 1.33× with
|
|
336
|
+
the collector enabled and 1.22× with it disabled (geometric means); without
|
|
337
|
+
the collector, `canada.json`, `mesh.json` and `numbers.json` are ties. yyjson 4.0.6 is left out: it
|
|
338
|
+
returns wrong strings for non-ASCII text.
|
|
339
|
+
|
|
340
|
+
Parsing is no longer the bottleneck. simdjson alone parses these files at
|
|
341
|
+
3.1 GB/s. It accounts for about a third of the time of `loads`; the rest goes
|
|
342
|
+
into creating Python objects. Freeing those objects later costs about a sixth
|
|
343
|
+
of the total. Even if parsing took no time at all, `loads` would be less than
|
|
344
|
+
1.5 times faster.
|
|
345
|
+
|
|
346
|
+
### Parts of documents with `parse`
|
|
347
|
+
|
|
348
|
+
If you only need a few values, `parse` creates only those. Extracting the id
|
|
349
|
+
and the screen name of the 100 statuses of `twitter.json`:
|
|
350
|
+
|
|
351
|
+
<picture>
|
|
352
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_parse_dark.png">
|
|
353
|
+
<img src="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_parse.png" width="80%" alt="Reading part of twitter.json: fastsimdjson parse 158 µs, pysimdjson 179, cysimdjson 229, msgspec 336, fastsimdjson loads 861, orjson 1009, json 3922">
|
|
354
|
+
</picture>
|
|
355
|
+
|
|
356
|
+
| method | µs |
|
|
357
|
+
|---|---:|
|
|
358
|
+
| `json.loads` | 3922 |
|
|
359
|
+
| orjson | 1009 |
|
|
360
|
+
| fastsimdjson `loads` | 861 |
|
|
361
|
+
| msgspec (typed `Struct`) | 336 |
|
|
362
|
+
| cysimdjson (lazy) | 229 |
|
|
363
|
+
| pysimdjson (lazy) | 179 |
|
|
364
|
+
| **fastsimdjson `parse`** | **158** |
|
|
365
|
+
|
|
366
|
+
Here `parse` is about 25 times faster than `json.loads` and 5.5 times faster
|
|
367
|
+
than `loads`. Most of its time is the simdjson parse itself.
|
|
368
|
+
|
|
369
|
+
### `dumps` against `json.dumps`
|
|
370
|
+
|
|
371
|
+
<picture>
|
|
372
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_dumps_dark.png">
|
|
373
|
+
<img src="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_dumps.png" width="80%" alt="Writing JSON as str: fastsimdjson dumps 4.63 times faster than json.dumps">
|
|
374
|
+
</picture>
|
|
375
|
+
|
|
376
|
+
`dumps` returns exactly the `str` of `json.dumps` and is 4.6 times faster
|
|
377
|
+
(geometric mean over the 22 files; from 3.1 times on `citm_catalog.json` to
|
|
378
|
+
12 times on `numbers.json`).
|
|
379
|
+
|
|
380
|
+
### `dumpb` against orjson and msgspec
|
|
381
|
+
|
|
382
|
+
`dumpb`, `orjson.dumps` and `msgspec.json.encode` produce compact UTF-8
|
|
383
|
+
`bytes`; `dumpb` and orjson produce identical bytes on every file. With the
|
|
384
|
+
standard library, the same compact bytes come from
|
|
385
|
+
`json.dumps(obj, separators=(",", ":"), ensure_ascii=False).encode()`.
|
|
386
|
+
Microseconds:
|
|
387
|
+
|
|
388
|
+
<picture>
|
|
389
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_dumpb_dark.png">
|
|
390
|
+
<img src="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_dumpb.png" width="80%" alt="Writing JSON as bytes, speed relative to json.dumps(...).encode(): fastsimdjson dumpb 10.09x, orjson 9.69x, msgspec 6.03x">
|
|
391
|
+
</picture>
|
|
392
|
+
|
|
393
|
+
| file | `json.dumps(...).encode()` | fastsimdjson `dumpb` | orjson | msgspec |
|
|
394
|
+
|---|---:|---:|---:|---:|
|
|
395
|
+
| twitter | 1827 | 202 | 204 | 330 |
|
|
396
|
+
| citm_catalog | 2921 | 432 | 432 | 499 |
|
|
397
|
+
| github_events | 177 | 17 | 19 | 31 |
|
|
398
|
+
| gsoc-2018 | 15386 | 595 | 554 | 1462 |
|
|
399
|
+
| update-center | 2482 | 254 | 232 | 477 |
|
|
400
|
+
| canada | 38101 | 2432 | 2936 | 3640 |
|
|
401
|
+
| mesh | 8726 | 785 | 999 | 1370 |
|
|
402
|
+
| numbers | 2530 | 178 | 200 | 332 |
|
|
403
|
+
|
|
404
|
+
`dumpb` and orjson are on par: over the 22 files, `dumpb` is 4% faster
|
|
405
|
+
(geometric mean), from 10% slower on text-heavy or tiny files
|
|
406
|
+
(`gsoc-2018.json`, `update-center.json`, `repeat.json`) to 27% faster on files
|
|
407
|
+
full of numbers. Both are 1.7
|
|
408
|
+
times faster than msgspec and 10 times faster than `json`. The comparison was
|
|
409
|
+
run on a processor with AVX-512, which `dumpb` and orjson both use to escape
|
|
410
|
+
strings; other x64 processors use SSE2 and ARM processors NEON.
|
|
411
|
+
|
|
412
|
+
### Streams with `loads_many`
|
|
413
|
+
|
|
414
|
+
20 MB of NDJSON (5268 objects and arrays, one per line), made from the same
|
|
415
|
+
files; best of three runs:
|
|
416
|
+
|
|
417
|
+
<picture>
|
|
418
|
+
<source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_ndjson_dark.png">
|
|
419
|
+
<img src="https://raw.githubusercontent.com/simdjson/fastpysimdjson/main/doc/perf_ndjson.png" width="80%" alt="Reading NDJSON: fastsimdjson parse_many 0.94 GB/s, loads_many 0.38, msgspec decode_lines 0.37, fastsimdjson loads per line 0.33, msgspec decode per line 0.27, orjson per line 0.26, json per line 0.12">
|
|
420
|
+
</picture>
|
|
421
|
+
|
|
422
|
+
| method | ms | GB/s |
|
|
423
|
+
|---|---:|---:|
|
|
424
|
+
| `json.loads` on each line | 169 | 0.12 |
|
|
425
|
+
| orjson on each line | 77 | 0.26 |
|
|
426
|
+
| msgspec `decode` on each line | 74 | 0.27 |
|
|
427
|
+
| fastsimdjson `loads` on each line | 60 | 0.33 |
|
|
428
|
+
| msgspec `decode_lines` | 53 | 0.37 |
|
|
429
|
+
| fastsimdjson `loads_many` | 52 | 0.38 |
|
|
430
|
+
| fastsimdjson `parse_many` (views only) | 21 | 0.94 |
|
|
431
|
+
|
|
432
|
+
msgspec's `decode_lines` and `loads_many` are on par (`loads_many` is about 2%
|
|
433
|
+
faster). `parse_many` only creates the views; reading values from them adds to
|
|
434
|
+
its time.
|
|
435
|
+
|
|
436
|
+
## Limitations
|
|
437
|
+
|
|
438
|
+
* The simdjson parser and the key and string caches are thread-local.
|
|
439
|
+
`release()` frees the parser and the cached strings retained by the calling
|
|
440
|
+
thread. A parser that grows past 64 MB is freed on its own at the end of
|
|
441
|
+
that call; its caches stay. A thread that exits without `release()` leaves
|
|
442
|
+
its cached strings behind.
|
|
443
|
+
* The module is marked free-threading compatible (`Py_MOD_GIL_NOT_USED` on
|
|
444
|
+
Python 3.13 and newer), so importing it on a free-threaded build does not
|
|
445
|
+
re-enable the GIL. On a free-threaded build, `bytearray` and `memoryview`
|
|
446
|
+
inputs are copied before parsing. Subinterpreters are not supported.
|