fastsimdjson 0.1.0__tar.gz → 0.2.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: fastsimdjson
3
- Version: 0.1.0
3
+ Version: 0.2.0
4
4
  Summary: Fast JSON parsing for Python, built on simdjson
5
5
  Author-email: Daniel Lemire <daniel@lemire.me>
6
6
  License-Expression: Apache-2.0
@@ -46,6 +46,36 @@ the exception is re-raised with Python's message and byte position. Any other
46
46
  exception from `json.loads` propagates. `release()` frees the simdjson parser
47
47
  and the string caches kept by the calling thread.
48
48
 
49
+ ## Lazy parsing
50
+
51
+ When you need only part of a document, `parse` avoids building the rest.
52
+ It accepts the same inputs as `loads` and returns read-only views:
53
+ `fastsimdjson.Object` (a `Mapping`) and `fastsimdjson.Array` (a
54
+ `Sequence`). Values are converted when you access them; nested objects and
55
+ arrays are returned as views. A scalar root is returned as a plain value.
56
+
57
+ ```python
58
+ doc = fastsimdjson.parse(open("twitter.json", "rb").read())
59
+ ids = [(s["id"], s["user"]["screen_name"]) for s in doc["statuses"]]
60
+ doc.at_pointer("/statuses/0/user/name") # JSON Pointer (RFC 6901)
61
+ doc["search_metadata"].as_dict() # convert a subtree, like loads
62
+ ```
63
+
64
+ `Object` supports `obj[key]`, `get`, `in`, `len`, iteration over the keys,
65
+ `keys()`, `values()`, `items()` (iterators), `at_pointer` and `as_dict()`.
66
+ `Array` supports `arr[i]` (negative indexes and slices), `len`, iteration,
67
+ `at_pointer` and `as_list()`. Both work with `match` statements.
68
+
69
+ * A view keeps its document alive; the document owns its own buffers, so
70
+ it remains valid while other documents are parsed.
71
+ * A key lookup scans the object. With duplicate keys, lookups return the
72
+ first value, whereas `as_dict()` (like `json.loads`) keeps the last.
73
+ * Indexing an array walks it from the last index reached, so a loop over
74
+ `arr[i]` is linear; iteration is the fastest way to visit an array.
75
+ * A document that simdjson rejects but `json.loads` accepts (an overflowing
76
+ number, an unpaired surrogate) is returned as plain Python objects, as
77
+ `loads` would return it.
78
+
49
79
  ## How it works
50
80
 
51
81
  1. simdjson's DOM parser (with runtime CPU dispatch: AVX-512, AVX2, SSE4.2, ...)
@@ -79,7 +109,7 @@ accepts it, so `loads` returns that string. A leading UTF-8 BOM is accepted.
79
109
 
80
110
  ## Build and test
81
111
 
82
- Python 3.10 or newer, and a C++17 compiler (clang or GCC). The simdjson 5.0.1
112
+ Python 3.10 or newer, and a C++17 compiler (clang, GCC or MSVC). The simdjson 5.0.2
83
113
  and simdutf 9.2.1 amalgamations are already in `vendor/`.
84
114
 
85
115
  pip:
@@ -18,6 +18,36 @@ the exception is re-raised with Python's message and byte position. Any other
18
18
  exception from `json.loads` propagates. `release()` frees the simdjson parser
19
19
  and the string caches kept by the calling thread.
20
20
 
21
+ ## Lazy parsing
22
+
23
+ When you need only part of a document, `parse` avoids building the rest.
24
+ It accepts the same inputs as `loads` and returns read-only views:
25
+ `fastsimdjson.Object` (a `Mapping`) and `fastsimdjson.Array` (a
26
+ `Sequence`). Values are converted when you access them; nested objects and
27
+ arrays are returned as views. A scalar root is returned as a plain value.
28
+
29
+ ```python
30
+ doc = fastsimdjson.parse(open("twitter.json", "rb").read())
31
+ ids = [(s["id"], s["user"]["screen_name"]) for s in doc["statuses"]]
32
+ doc.at_pointer("/statuses/0/user/name") # JSON Pointer (RFC 6901)
33
+ doc["search_metadata"].as_dict() # convert a subtree, like loads
34
+ ```
35
+
36
+ `Object` supports `obj[key]`, `get`, `in`, `len`, iteration over the keys,
37
+ `keys()`, `values()`, `items()` (iterators), `at_pointer` and `as_dict()`.
38
+ `Array` supports `arr[i]` (negative indexes and slices), `len`, iteration,
39
+ `at_pointer` and `as_list()`. Both work with `match` statements.
40
+
41
+ * A view keeps its document alive; the document owns its own buffers, so
42
+ it remains valid while other documents are parsed.
43
+ * A key lookup scans the object. With duplicate keys, lookups return the
44
+ first value, whereas `as_dict()` (like `json.loads`) keeps the last.
45
+ * Indexing an array walks it from the last index reached, so a loop over
46
+ `arr[i]` is linear; iteration is the fastest way to visit an array.
47
+ * A document that simdjson rejects but `json.loads` accepts (an overflowing
48
+ number, an unpaired surrogate) is returned as plain Python objects, as
49
+ `loads` would return it.
50
+
21
51
  ## How it works
22
52
 
23
53
  1. simdjson's DOM parser (with runtime CPU dispatch: AVX-512, AVX2, SSE4.2, ...)
@@ -51,7 +81,7 @@ accepts it, so `loads` returns that string. A leading UTF-8 BOM is accepted.
51
81
 
52
82
  ## Build and test
53
83
 
54
- Python 3.10 or newer, and a C++17 compiler (clang or GCC). The simdjson 5.0.1
84
+ Python 3.10 or newer, and a C++17 compiler (clang, GCC or MSVC). The simdjson 5.0.2
55
85
  and simdutf 9.2.1 amalgamations are already in `vendor/`.
56
86
 
57
87
  pip:
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "fastsimdjson"
7
- version = "0.1.0"
7
+ version = "0.2.0"
8
8
  description = "Fast JSON parsing for Python, built on simdjson"
9
9
  readme = "README.md"
10
10
  authors = [