pydoptic 0.0.1.post1.dev2__tar.gz → 0.0.2.post1.dev1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {pydoptic-0.0.1.post1.dev2/src/pydoptic.egg-info → pydoptic-0.0.2.post1.dev1}/PKG-INFO +6 -229
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/README.md +5 -228
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1/src/pydoptic.egg-info}/PKG-INFO +6 -229
- pydoptic-0.0.2.post1.dev1/src/pydoptic.egg-info/scm_version.json +8 -0
- pydoptic-0.0.1.post1.dev2/src/pydoptic.egg-info/scm_version.json +0 -8
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/LICENSE +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/pyproject.toml +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/setup.cfg +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic/__init__.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic/base_model.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic/py.typed +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic/selector.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic/validate_types.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/SOURCES.txt +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/dependency_links.txt +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/requires.txt +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/scm_file_list.json +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/top_level.txt +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/test/test_base_model.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/test/test_partial_model.py +0 -0
- {pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/test/test_selector.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: pydoptic
|
|
3
|
-
Version: 0.0.
|
|
3
|
+
Version: 0.0.2.post1.dev1
|
|
4
4
|
Summary: Model data using reified optics
|
|
5
5
|
Author: John Hungerford
|
|
6
6
|
License-Expression: MIT
|
|
@@ -16,14 +16,12 @@ Provides-Extra: benchmark
|
|
|
16
16
|
Requires-Dist: pydantic~=2.11; extra == "benchmark"
|
|
17
17
|
Dynamic: license-file
|
|
18
18
|
|
|
19
|
-
#
|
|
19
|
+
# pydoptic
|
|
20
20
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
2. You frequently need to retrieve and manipulate deeply nested values
|
|
26
|
-
3. You need a type-safe way to refer to properties when querying remote data
|
|
21
|
+
The core data modeling library based on reified optics. See the
|
|
22
|
+
[repository README](https://github.com/johnhungerford/pydoptic) for what reified optics are and why
|
|
23
|
+
they're useful -- this document is the in-depth user guide: defining and constructing models, reading
|
|
24
|
+
and mutating data, chaining selects, discriminated subtypes, and integrating with external APIs.
|
|
27
25
|
|
|
28
26
|
## Installation
|
|
29
27
|
|
|
@@ -31,227 +29,6 @@ A good alternative to Pydantic when one or more of the following is true:
|
|
|
31
29
|
pip install pydoptic
|
|
32
30
|
```
|
|
33
31
|
|
|
34
|
-
This is the core data modeling library only. For the SQL query builder or the Elasticsearch
|
|
35
|
-
integration built on top of it, see [pydoptic-sql](https://pypi.org/project/pydoptic-sql/) and
|
|
36
|
-
[pydoptic-elastic](https://pypi.org/project/pydoptic-elastic/) -- both separate packages, published
|
|
37
|
-
from the same [repository](https://github.com/johnhungerford/pydoptic).
|
|
38
|
-
|
|
39
|
-
## Why?
|
|
40
|
-
|
|
41
|
-
In every programming language, data is modeled as the type of container used to store it in memory. In Python, this generally looks something like the following:
|
|
42
|
-
|
|
43
|
-
```python3
|
|
44
|
-
from pydantic import BaseModel
|
|
45
|
-
from datetime import date
|
|
46
|
-
from typing import List
|
|
47
|
-
|
|
48
|
-
class Address(BaseModel):
|
|
49
|
-
number: str
|
|
50
|
-
street: str
|
|
51
|
-
city: str
|
|
52
|
-
postal_code: str
|
|
53
|
-
state: str
|
|
54
|
-
|
|
55
|
-
class Organization(BaseModel):
|
|
56
|
-
name: str
|
|
57
|
-
address: Address
|
|
58
|
-
phone_number: str
|
|
59
|
-
owner: 'Person'
|
|
60
|
-
members: List['Person'] | None
|
|
61
|
-
|
|
62
|
-
class Person(BaseModel):
|
|
63
|
-
name: str
|
|
64
|
-
address: Address
|
|
65
|
-
birth_date: date
|
|
66
|
-
phone_number: str
|
|
67
|
-
organizations: List[Organization] | None
|
|
68
|
-
is_active: bool
|
|
69
|
-
```
|
|
70
|
-
|
|
71
|
-
While the above class definitions do provide a typed description of the data we expect to deal with, their utility when it comes to *doing* anything with the data is actually fairly limited. At most they can do the following:
|
|
72
|
-
|
|
73
|
-
1. Validate and then represent a complete instance of each data type, either from parameters or some serialized source (e.g., JSON)
|
|
74
|
-
2. Provide access to any property that can be validated with a type checker
|
|
75
|
-
3. Generate a valid serialized representation of an instance (e.g. JSON)
|
|
76
|
-
|
|
77
|
-
While this is certainly useful, consider all the other things we might want to do:
|
|
78
|
-
|
|
79
|
-
1. Consume and validate an *incomplete* instance of our data (e.g., a `Person`'s `name` and `birth_date`, but nothing else) without making the *complete* data model less precise (e.g., by making `Person.address` optional)
|
|
80
|
-
2. Access potentially missing values of an incomplete model
|
|
81
|
-
3. Query just the `name` and `birth_date` from a database. Nothing in the data model allows us to *specify the field itself* -- only the *value* of the field. This is an important distinction.
|
|
82
|
-
4. Update the `is_active` flag of every member of every organization a given person is a member of without having to check if `Person.organizations` or `Organization.members` are `None`.
|
|
83
|
-
|
|
84
|
-
Achieving these things with ordinary data models is inconvenient at best. Dealing with incomplete data requires defining new types for each use case or abandoning type safety altogether. While data models can read and write (complete) records to a datastore, they provide no mechanism for helping us specify the parts of the data model that are of interest to us, like when we want to specify which fields to retrieve or define a query setting some constraint on a particular field. In general we have to specify field names as strings, losing any pre-runtime validation. Finally, traditional data models provide no abstractions for manipulating nested data, nor do they provide any mechanism that could be used to create such abstractions.
|
|
85
|
-
|
|
86
|
-
Pydoptic provides an alternative way to model data in Python -- and indeed in any language -- that allows you to accomplish all these and things much more easily and safely.
|
|
87
|
-
|
|
88
|
-
## How?
|
|
89
|
-
|
|
90
|
-
Whereas traditional data models represent data as *containers* of properties, Pydoptic models data as collections of *references* to properties. In a traditional data type, a "property" is a value contained by an in-memory instance of the type in question. In Pydoptic, a "property" is a *description* of a value belonging to a type. This description can be used to do the usual things -- storing a value in an object or retrieving it from an object -- but it can also do a much more. For instance, it can be used to store a value in a *remote* data source, or query it and retrieve it from that data source, or... pretty much anything else you can think of that might need to be done with it!
|
|
91
|
-
|
|
92
|
-
Here is what a Pydoptic version of the data model above looks like:
|
|
93
|
-
|
|
94
|
-
```python3
|
|
95
|
-
from pydoptic import BaseModel, PartialModel, Prop, PropOpt, PropOptArr
|
|
96
|
-
from datetime import date
|
|
97
|
-
|
|
98
|
-
class Address(BaseModel):
|
|
99
|
-
number: Prop['Address', str]
|
|
100
|
-
street: Prop['Address', str]
|
|
101
|
-
city: Prop['Address', str]
|
|
102
|
-
postal_code: Prop['Address', str]
|
|
103
|
-
state: Prop['Address', str]
|
|
104
|
-
|
|
105
|
-
class Organization(BaseModel):
|
|
106
|
-
name: Prop['Organization', str]
|
|
107
|
-
address: Prop['Organization', Address]
|
|
108
|
-
phone_number: Prop['Organization', str]
|
|
109
|
-
owner: PropOpt['Organization', 'Person']
|
|
110
|
-
members: PropOptArr['Organization', 'Person']
|
|
111
|
-
|
|
112
|
-
class Person(BaseModel):
|
|
113
|
-
name: Prop['Person', str]
|
|
114
|
-
address: Prop['Person', Address]
|
|
115
|
-
birth_date: Prop['Person', date]
|
|
116
|
-
phone_number: Prop['Person', str]
|
|
117
|
-
organizations: PropOptArr['Person', Organization]
|
|
118
|
-
is_active: Prop['Person', bool]
|
|
119
|
-
```
|
|
120
|
-
|
|
121
|
-
The type signatures of the properties include some more boilerplate, to be sure. Most notably, they all contain references to the model class that they belong to. This may seem redundant, but it's a crucial feature that makes them as powerful as the are: because they contain references back to the classes they belong to, they can be used entirely independently of those classes. Each property -- which is only a *class* attribute -- is initialized with a value containing references to both its "origin" type (the model class) and its "target" type (the second type parameter on the right) along with its attribute name and flags indicating whether it's optional (for `PropOpt` properties), array (`PropArr`), or both (`PropOptArr`).
|
|
122
|
-
|
|
123
|
-
By constructing our properties as comprehensive metadata *about* values and their relationship with the model they belong to, rather than simply the values themselves, we provide ourselves with a much more flexible and powerful tool. Let's see what we can do with them.
|
|
124
|
-
|
|
125
|
-
### Basics
|
|
126
|
-
|
|
127
|
-
Let's start with the basics. These "props" would not be much use to us if we could not actually construct model instances. It turns out we can do this in the usual way:
|
|
128
|
-
|
|
129
|
-
```python3
|
|
130
|
-
person = Person(name="John", address=Address(...), birth_date=..., phone_number=..., is_active=True)
|
|
131
|
-
|
|
132
|
-
print(person.name)
|
|
133
|
-
# John
|
|
134
|
-
```
|
|
135
|
-
|
|
136
|
-
On initialization, the model class constructs the `Prop` class attributes based on the type hints and keeps track of them internally. It then uses the known properties to validate keyword arguments that are provided when constructing a class instance. The above succeeds even though `organizations` is missing because organizations is a `PropOptArr`, which is optional. If we left out `name`, however, it would raise a `ValueError`.
|
|
137
|
-
|
|
138
|
-
Note that your type checker will complain that `person.name` is a `Prop` instead of a `str`. While Pydoptic does assign properties as model attributes (the `Prop` types should be defined only on the class), your type checker will not know this. The "Pydoptic" way to retrieve properties is not to access the attribute directly, but use the property itself!
|
|
139
|
-
|
|
140
|
-
```python3
|
|
141
|
-
person_name = Person.name.get_val(person)
|
|
142
|
-
|
|
143
|
-
print(person_name)
|
|
144
|
-
# John
|
|
145
|
-
```
|
|
146
|
-
|
|
147
|
-
Your type checker will recognize `person_name` as having a type `str`. While this is a fairly verbose way of getting a simple value, its utility will become clearer when you find yourself frequently accessing nested data. For instance, say you want to flip the `is_active` status of every members of every organization a given person is connected with. Ordinarily this would require a fairly elaborate combination of `for` loops and and `if` statements:
|
|
148
|
-
|
|
149
|
-
```python3
|
|
150
|
-
if person.organizations is not None:
|
|
151
|
-
for organization in organizations:
|
|
152
|
-
if organization.members is not None:
|
|
153
|
-
for member in members:
|
|
154
|
-
member.is_active = !member.is_active
|
|
155
|
-
```
|
|
156
|
-
|
|
157
|
-
In this case, Pydoptic's model is substantially *less* verbose. We can reproduce all of the above logic by *composing* our `Prop`s to select the desired path to `is_active` and then use the `update` method to change every value:
|
|
158
|
-
|
|
159
|
-
```python3
|
|
160
|
-
select_related_statuses = Person.organizations(Organization.members)(Person.is_active)
|
|
161
|
-
|
|
162
|
-
statuses = select_related_statuses.get(person).as_list
|
|
163
|
-
print(statuses)
|
|
164
|
-
# [True, False, False, True, ...]
|
|
165
|
-
|
|
166
|
-
# All the updating is done here:
|
|
167
|
-
select_related_statuses.update(person, lambda status: !status)
|
|
168
|
-
|
|
169
|
-
statuses = select_related_statuses.get(person).as_list
|
|
170
|
-
print(statuses)
|
|
171
|
-
# [False, True, True, False, ...]
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
As your type checker should indicate, `select_related_statuses` is a `Select[Person, bool]` which is composed from the `Prop`s `Person.organizations`, `Organization.members`, and `Person.is_active`. This particular chain of props can compose because each prop's *target* type is the same as the *origin* type of the prop chained to it. When this is not the case (meaning the chaining is *invalid*), the type checker will complain. The resulting `Select` can then be used to retrieve and update all `Person.is_active` properties (the final `Prop` in the chain) nested within the original `person` via the path `organizations` -> `members` -> `is_active`. Since `Person.organizations` and `Organization.members` are optional properties, ordinarily retrieving and updating these values would require checking for `None` in multiple places.
|
|
175
|
-
|
|
176
|
-
### Incomplete data
|
|
177
|
-
|
|
178
|
-
By separating our property types from the actual container, handling incomplete data becomes simpler without losing type precision. The same `Prop`s we use to set and retrieve data a model like `Person` can be used to do the same in incomplete models as well.
|
|
179
|
-
|
|
180
|
-
```python3
|
|
181
|
-
partial_person: PartialModel[Person] = Person.partial(name="John", birth_date=...)
|
|
182
|
-
|
|
183
|
-
print(Person.name.get_val_unsafe(partial_person))
|
|
184
|
-
# John
|
|
185
|
-
|
|
186
|
-
person_dict = {'name': 'John'}
|
|
187
|
-
print(Person.name.get_val_unsafe(person_dict))
|
|
188
|
-
# John
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
As you can see in the example above, there are two ways to represent incomplete data. `PartialModel` is a model type that can have missing properties and extraneous properties, but any property whose name corresponds to a property on the full model must be valid. Hence `Person.partial(name=23)` would fail because `23` is not a string. This is useful when you are consuming partial data from a data source but still want to validate it. If you don't want to validate the data at all, you can also use a `dict`. If you know your data source is producing valid data, it probably makes sense to just keep your data in an untyped `Dict[str, Any]`.
|
|
192
|
-
|
|
193
|
-
All of the data manipulation methods used on full models have versions that can be used on `PartialModel`s and `dict`s. Methods with the suffix `_unsafe` will work exactly like the regular method, but will raise a `ValueError` when required data is missing (or invalid in certain ways); methods with the suffix `_safe` will return `None` when data is invalid or (for mutating methods) fail silently.
|
|
194
|
-
|
|
195
|
-
### Data integration
|
|
196
|
-
|
|
197
|
-
Since `Prop`s contain all the information about the properties they reference, they can easily be used for interacting with *remote* instances of the same data. This repository includes an example Elasticsearch integration as a reference. Here is what it looks like to use the pydoptic-based Elasticsearch API:
|
|
198
|
-
|
|
199
|
-
```python3
|
|
200
|
-
from pydoptic import Prop, PropOptArr, PartialModel
|
|
201
|
-
from pydoptic_elastic import ElasticModel, Query, ElasticService, elastic_prop, ESMapping
|
|
202
|
-
from datetime import date
|
|
203
|
-
from typing import List
|
|
204
|
-
from elasticsearch import Elasticsearch
|
|
205
|
-
|
|
206
|
-
class Address(ElasticModel):
|
|
207
|
-
...
|
|
208
|
-
|
|
209
|
-
class Organization(ElasticModel):
|
|
210
|
-
...
|
|
211
|
-
|
|
212
|
-
class Person(ElasticModel):
|
|
213
|
-
name: Prop['Person', str]
|
|
214
|
-
address: Prop['Person', Address]
|
|
215
|
-
birth_date: Prop['Person', date]
|
|
216
|
-
phone_number: Prop['Person', str] = elastic_prop(mapping=ESMapping.keyword)
|
|
217
|
-
organizations: PropOptArr['Person', Organization]
|
|
218
|
-
is_active: Prop['Person', bool]
|
|
219
|
-
|
|
220
|
-
elastic_service = ElasticService(Elasticsearch('http://localhost:9200'))
|
|
221
|
-
|
|
222
|
-
elastic_service.create_index(Person)
|
|
223
|
-
|
|
224
|
-
original_person = Person(name='John', ...)
|
|
225
|
-
|
|
226
|
-
elastic_service.index(original_person)
|
|
227
|
-
|
|
228
|
-
query: Query[Person] = Query.match(Person.name, Person.get_val(original_person))
|
|
229
|
-
|
|
230
|
-
found_people: List[PartialModel[Person]] = elastic_service.search_partial(query, source=[Person.name, Person.birth_date])
|
|
231
|
-
|
|
232
|
-
for found_person in found_people:
|
|
233
|
-
print(Person.name.get_val_unsafe(person))
|
|
234
|
-
# John
|
|
235
|
-
```
|
|
236
|
-
|
|
237
|
-
The above data model is built using the `ElasticModel` base class, which is a specialized subtype of `BaseModel` that captures property metadata specifically for Elasticsearch (e.g., the index name and field mappings). `Prop`s can be customized with `elastic_prop` to include field-level metadata like `mapping` (the `Prop` type has a metadata field for storing arbitrary key-value data for use cases like this).
|
|
238
|
-
|
|
239
|
-
`Query` is a representation of queries based on Pydoptic types. For instance, the match query is constructed by using a `Prop` value to specify the field to match with and providing a value corresponding to that prop's target type. The result, when querying `Person.name` is a `Query[Person]`. `Query[Person]` contains a reference to the `Person` class, which can be used to resolve the appropriate index name.
|
|
240
|
-
|
|
241
|
-
`ElasticService` provides an API for dealing with indices and documents using the Pydoptic-based `ElasticModel` and `Query` types. We first create our `Person` index by simply passing the class to `elastic_service.create_index`. The index name is generated by default from the class name (`person`) and the field mappings are constructed from the properties. We then index a document by passing it to `elastic_service.index`; since the instance contains a reference to the class, the index name can be resolved properly. Finally, we search for the original person by using our `Query[Person]`, which matches on the original person's name. When searching, however, we use a special variant `elastic_service.search_partial` which allows us to provide a `source` parameter, where we specify only two `Person` properties: `Person.name` and `Person.birth_date`. The result is a list of `PartialModel[Person]` instances containing only those fields.
|
|
242
|
-
|
|
243
|
-
You'll notice at no point in the above are we forced to pass any index or property names as strings. All the information required to resolve the indices and fields are contained in our model types.
|
|
244
|
-
|
|
245
|
-
### Reified optics
|
|
246
|
-
|
|
247
|
-
I hope the above has indicated clearly enough how pydoptic can be used to solve the problems indicated in the first section. At this point its worth saying a few words about the approach used.
|
|
248
|
-
|
|
249
|
-
Pydoptic is inspired by a concept from functional programming called "optics". In functional programming, optics are not just useful but pretty much necessary due to the relative difficulty of updating nested properties in immutable data structures. Since updating nested data is easier in imperative languages like Python, optics tend not to be used much. There a couple of optics libraries out there for Python, but they try to be functional in the full sense, which is to say they are designed to create immutable copies of dataclasses rather than mutate them.
|
|
250
|
-
|
|
251
|
-
Pydoptic's approach differs from most optics libraries in two ways. First, it takes the compositional properties of optic types from functional programming while giving them power to mutate objects. Second, it "reifies" the optics. Functional optics are typically encoded as *functions* that retrieve or update data and can be composed in various ways; in Pydoptic, they are encoded as *data*. They are *descriptions* of the things that *could be* accessed or updated. The actual functionality for doing the accessing/updating is secondary, and is implemented by *interpreting* the descriptions. This *reification* of the optics gives it a more general power than traditional optics.
|
|
252
|
-
|
|
253
|
-
This concept of reified optics comes from the Scala project [ZIO schema](https://zio.dev/zio-schema/), which provides (among other things) a similar mechanism for referencing properties via "accessors". Pydoptic provides a simplified version of this approach appropriate to Python and its more limited (though still quite powerful!) type system. The main innovation of Pydoptic (beyond bringing optics to mutable data) is that it models data "optics-first," unlike ZIO-schema, which (for good reasons connected with Scala and the JVM) *derives* optics *from* traditional data models (i.e., data classes).
|
|
254
|
-
|
|
255
32
|
## User guide
|
|
256
33
|
|
|
257
34
|
The examples below assume this import, unless a different one is shown:
|
|
@@ -1,11 +1,9 @@
|
|
|
1
|
-
#
|
|
1
|
+
# pydoptic
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
2. You frequently need to retrieve and manipulate deeply nested values
|
|
8
|
-
3. You need a type-safe way to refer to properties when querying remote data
|
|
3
|
+
The core data modeling library based on reified optics. See the
|
|
4
|
+
[repository README](https://github.com/johnhungerford/pydoptic) for what reified optics are and why
|
|
5
|
+
they're useful -- this document is the in-depth user guide: defining and constructing models, reading
|
|
6
|
+
and mutating data, chaining selects, discriminated subtypes, and integrating with external APIs.
|
|
9
7
|
|
|
10
8
|
## Installation
|
|
11
9
|
|
|
@@ -13,227 +11,6 @@ A good alternative to Pydantic when one or more of the following is true:
|
|
|
13
11
|
pip install pydoptic
|
|
14
12
|
```
|
|
15
13
|
|
|
16
|
-
This is the core data modeling library only. For the SQL query builder or the Elasticsearch
|
|
17
|
-
integration built on top of it, see [pydoptic-sql](https://pypi.org/project/pydoptic-sql/) and
|
|
18
|
-
[pydoptic-elastic](https://pypi.org/project/pydoptic-elastic/) -- both separate packages, published
|
|
19
|
-
from the same [repository](https://github.com/johnhungerford/pydoptic).
|
|
20
|
-
|
|
21
|
-
## Why?
|
|
22
|
-
|
|
23
|
-
In every programming language, data is modeled as the type of container used to store it in memory. In Python, this generally looks something like the following:
|
|
24
|
-
|
|
25
|
-
```python3
|
|
26
|
-
from pydantic import BaseModel
|
|
27
|
-
from datetime import date
|
|
28
|
-
from typing import List
|
|
29
|
-
|
|
30
|
-
class Address(BaseModel):
|
|
31
|
-
number: str
|
|
32
|
-
street: str
|
|
33
|
-
city: str
|
|
34
|
-
postal_code: str
|
|
35
|
-
state: str
|
|
36
|
-
|
|
37
|
-
class Organization(BaseModel):
|
|
38
|
-
name: str
|
|
39
|
-
address: Address
|
|
40
|
-
phone_number: str
|
|
41
|
-
owner: 'Person'
|
|
42
|
-
members: List['Person'] | None
|
|
43
|
-
|
|
44
|
-
class Person(BaseModel):
|
|
45
|
-
name: str
|
|
46
|
-
address: Address
|
|
47
|
-
birth_date: date
|
|
48
|
-
phone_number: str
|
|
49
|
-
organizations: List[Organization] | None
|
|
50
|
-
is_active: bool
|
|
51
|
-
```
|
|
52
|
-
|
|
53
|
-
While the above class definitions do provide a typed description of the data we expect to deal with, their utility when it comes to *doing* anything with the data is actually fairly limited. At most they can do the following:
|
|
54
|
-
|
|
55
|
-
1. Validate and then represent a complete instance of each data type, either from parameters or some serialized source (e.g., JSON)
|
|
56
|
-
2. Provide access to any property that can be validated with a type checker
|
|
57
|
-
3. Generate a valid serialized representation of an instance (e.g. JSON)
|
|
58
|
-
|
|
59
|
-
While this is certainly useful, consider all the other things we might want to do:
|
|
60
|
-
|
|
61
|
-
1. Consume and validate an *incomplete* instance of our data (e.g., a `Person`'s `name` and `birth_date`, but nothing else) without making the *complete* data model less precise (e.g., by making `Person.address` optional)
|
|
62
|
-
2. Access potentially missing values of an incomplete model
|
|
63
|
-
3. Query just the `name` and `birth_date` from a database. Nothing in the data model allows us to *specify the field itself* -- only the *value* of the field. This is an important distinction.
|
|
64
|
-
4. Update the `is_active` flag of every member of every organization a given person is a member of without having to check if `Person.organizations` or `Organization.members` are `None`.
|
|
65
|
-
|
|
66
|
-
Achieving these things with ordinary data models is inconvenient at best. Dealing with incomplete data requires defining new types for each use case or abandoning type safety altogether. While data models can read and write (complete) records to a datastore, they provide no mechanism for helping us specify the parts of the data model that are of interest to us, like when we want to specify which fields to retrieve or define a query setting some constraint on a particular field. In general we have to specify field names as strings, losing any pre-runtime validation. Finally, traditional data models provide no abstractions for manipulating nested data, nor do they provide any mechanism that could be used to create such abstractions.
|
|
67
|
-
|
|
68
|
-
Pydoptic provides an alternative way to model data in Python -- and indeed in any language -- that allows you to accomplish all these and things much more easily and safely.
|
|
69
|
-
|
|
70
|
-
## How?
|
|
71
|
-
|
|
72
|
-
Whereas traditional data models represent data as *containers* of properties, Pydoptic models data as collections of *references* to properties. In a traditional data type, a "property" is a value contained by an in-memory instance of the type in question. In Pydoptic, a "property" is a *description* of a value belonging to a type. This description can be used to do the usual things -- storing a value in an object or retrieving it from an object -- but it can also do a much more. For instance, it can be used to store a value in a *remote* data source, or query it and retrieve it from that data source, or... pretty much anything else you can think of that might need to be done with it!
|
|
73
|
-
|
|
74
|
-
Here is what a Pydoptic version of the data model above looks like:
|
|
75
|
-
|
|
76
|
-
```python3
|
|
77
|
-
from pydoptic import BaseModel, PartialModel, Prop, PropOpt, PropOptArr
|
|
78
|
-
from datetime import date
|
|
79
|
-
|
|
80
|
-
class Address(BaseModel):
|
|
81
|
-
number: Prop['Address', str]
|
|
82
|
-
street: Prop['Address', str]
|
|
83
|
-
city: Prop['Address', str]
|
|
84
|
-
postal_code: Prop['Address', str]
|
|
85
|
-
state: Prop['Address', str]
|
|
86
|
-
|
|
87
|
-
class Organization(BaseModel):
|
|
88
|
-
name: Prop['Organization', str]
|
|
89
|
-
address: Prop['Organization', Address]
|
|
90
|
-
phone_number: Prop['Organization', str]
|
|
91
|
-
owner: PropOpt['Organization', 'Person']
|
|
92
|
-
members: PropOptArr['Organization', 'Person']
|
|
93
|
-
|
|
94
|
-
class Person(BaseModel):
|
|
95
|
-
name: Prop['Person', str]
|
|
96
|
-
address: Prop['Person', Address]
|
|
97
|
-
birth_date: Prop['Person', date]
|
|
98
|
-
phone_number: Prop['Person', str]
|
|
99
|
-
organizations: PropOptArr['Person', Organization]
|
|
100
|
-
is_active: Prop['Person', bool]
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
The type signatures of the properties include some more boilerplate, to be sure. Most notably, they all contain references to the model class that they belong to. This may seem redundant, but it's a crucial feature that makes them as powerful as the are: because they contain references back to the classes they belong to, they can be used entirely independently of those classes. Each property -- which is only a *class* attribute -- is initialized with a value containing references to both its "origin" type (the model class) and its "target" type (the second type parameter on the right) along with its attribute name and flags indicating whether it's optional (for `PropOpt` properties), array (`PropArr`), or both (`PropOptArr`).
|
|
104
|
-
|
|
105
|
-
By constructing our properties as comprehensive metadata *about* values and their relationship with the model they belong to, rather than simply the values themselves, we provide ourselves with a much more flexible and powerful tool. Let's see what we can do with them.
|
|
106
|
-
|
|
107
|
-
### Basics
|
|
108
|
-
|
|
109
|
-
Let's start with the basics. These "props" would not be much use to us if we could not actually construct model instances. It turns out we can do this in the usual way:
|
|
110
|
-
|
|
111
|
-
```python3
|
|
112
|
-
person = Person(name="John", address=Address(...), birth_date=..., phone_number=..., is_active=True)
|
|
113
|
-
|
|
114
|
-
print(person.name)
|
|
115
|
-
# John
|
|
116
|
-
```
|
|
117
|
-
|
|
118
|
-
On initialization, the model class constructs the `Prop` class attributes based on the type hints and keeps track of them internally. It then uses the known properties to validate keyword arguments that are provided when constructing a class instance. The above succeeds even though `organizations` is missing because organizations is a `PropOptArr`, which is optional. If we left out `name`, however, it would raise a `ValueError`.
|
|
119
|
-
|
|
120
|
-
Note that your type checker will complain that `person.name` is a `Prop` instead of a `str`. While Pydoptic does assign properties as model attributes (the `Prop` types should be defined only on the class), your type checker will not know this. The "Pydoptic" way to retrieve properties is not to access the attribute directly, but use the property itself!
|
|
121
|
-
|
|
122
|
-
```python3
|
|
123
|
-
person_name = Person.name.get_val(person)
|
|
124
|
-
|
|
125
|
-
print(person_name)
|
|
126
|
-
# John
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
Your type checker will recognize `person_name` as having a type `str`. While this is a fairly verbose way of getting a simple value, its utility will become clearer when you find yourself frequently accessing nested data. For instance, say you want to flip the `is_active` status of every members of every organization a given person is connected with. Ordinarily this would require a fairly elaborate combination of `for` loops and and `if` statements:
|
|
130
|
-
|
|
131
|
-
```python3
|
|
132
|
-
if person.organizations is not None:
|
|
133
|
-
for organization in organizations:
|
|
134
|
-
if organization.members is not None:
|
|
135
|
-
for member in members:
|
|
136
|
-
member.is_active = !member.is_active
|
|
137
|
-
```
|
|
138
|
-
|
|
139
|
-
In this case, Pydoptic's model is substantially *less* verbose. We can reproduce all of the above logic by *composing* our `Prop`s to select the desired path to `is_active` and then use the `update` method to change every value:
|
|
140
|
-
|
|
141
|
-
```python3
|
|
142
|
-
select_related_statuses = Person.organizations(Organization.members)(Person.is_active)
|
|
143
|
-
|
|
144
|
-
statuses = select_related_statuses.get(person).as_list
|
|
145
|
-
print(statuses)
|
|
146
|
-
# [True, False, False, True, ...]
|
|
147
|
-
|
|
148
|
-
# All the updating is done here:
|
|
149
|
-
select_related_statuses.update(person, lambda status: !status)
|
|
150
|
-
|
|
151
|
-
statuses = select_related_statuses.get(person).as_list
|
|
152
|
-
print(statuses)
|
|
153
|
-
# [False, True, True, False, ...]
|
|
154
|
-
```
|
|
155
|
-
|
|
156
|
-
As your type checker should indicate, `select_related_statuses` is a `Select[Person, bool]` which is composed from the `Prop`s `Person.organizations`, `Organization.members`, and `Person.is_active`. This particular chain of props can compose because each prop's *target* type is the same as the *origin* type of the prop chained to it. When this is not the case (meaning the chaining is *invalid*), the type checker will complain. The resulting `Select` can then be used to retrieve and update all `Person.is_active` properties (the final `Prop` in the chain) nested within the original `person` via the path `organizations` -> `members` -> `is_active`. Since `Person.organizations` and `Organization.members` are optional properties, ordinarily retrieving and updating these values would require checking for `None` in multiple places.
|
|
157
|
-
|
|
158
|
-
### Incomplete data
|
|
159
|
-
|
|
160
|
-
By separating our property types from the actual container, handling incomplete data becomes simpler without losing type precision. The same `Prop`s we use to set and retrieve data a model like `Person` can be used to do the same in incomplete models as well.
|
|
161
|
-
|
|
162
|
-
```python3
|
|
163
|
-
partial_person: PartialModel[Person] = Person.partial(name="John", birth_date=...)
|
|
164
|
-
|
|
165
|
-
print(Person.name.get_val_unsafe(partial_person))
|
|
166
|
-
# John
|
|
167
|
-
|
|
168
|
-
person_dict = {'name': 'John'}
|
|
169
|
-
print(Person.name.get_val_unsafe(person_dict))
|
|
170
|
-
# John
|
|
171
|
-
```
|
|
172
|
-
|
|
173
|
-
As you can see in the example above, there are two ways to represent incomplete data. `PartialModel` is a model type that can have missing properties and extraneous properties, but any property whose name corresponds to a property on the full model must be valid. Hence `Person.partial(name=23)` would fail because `23` is not a string. This is useful when you are consuming partial data from a data source but still want to validate it. If you don't want to validate the data at all, you can also use a `dict`. If you know your data source is producing valid data, it probably makes sense to just keep your data in an untyped `Dict[str, Any]`.
|
|
174
|
-
|
|
175
|
-
All of the data manipulation methods used on full models have versions that can be used on `PartialModel`s and `dict`s. Methods with the suffix `_unsafe` will work exactly like the regular method, but will raise a `ValueError` when required data is missing (or invalid in certain ways); methods with the suffix `_safe` will return `None` when data is invalid or (for mutating methods) fail silently.
|
|
176
|
-
|
|
177
|
-
### Data integration
|
|
178
|
-
|
|
179
|
-
Since `Prop`s contain all the information about the properties they reference, they can easily be used for interacting with *remote* instances of the same data. This repository includes an example Elasticsearch integration as a reference. Here is what it looks like to use the pydoptic-based Elasticsearch API:
|
|
180
|
-
|
|
181
|
-
```python3
|
|
182
|
-
from pydoptic import Prop, PropOptArr, PartialModel
|
|
183
|
-
from pydoptic_elastic import ElasticModel, Query, ElasticService, elastic_prop, ESMapping
|
|
184
|
-
from datetime import date
|
|
185
|
-
from typing import List
|
|
186
|
-
from elasticsearch import Elasticsearch
|
|
187
|
-
|
|
188
|
-
class Address(ElasticModel):
|
|
189
|
-
...
|
|
190
|
-
|
|
191
|
-
class Organization(ElasticModel):
|
|
192
|
-
...
|
|
193
|
-
|
|
194
|
-
class Person(ElasticModel):
|
|
195
|
-
name: Prop['Person', str]
|
|
196
|
-
address: Prop['Person', Address]
|
|
197
|
-
birth_date: Prop['Person', date]
|
|
198
|
-
phone_number: Prop['Person', str] = elastic_prop(mapping=ESMapping.keyword)
|
|
199
|
-
organizations: PropOptArr['Person', Organization]
|
|
200
|
-
is_active: Prop['Person', bool]
|
|
201
|
-
|
|
202
|
-
elastic_service = ElasticService(Elasticsearch('http://localhost:9200'))
|
|
203
|
-
|
|
204
|
-
elastic_service.create_index(Person)
|
|
205
|
-
|
|
206
|
-
original_person = Person(name='John', ...)
|
|
207
|
-
|
|
208
|
-
elastic_service.index(original_person)
|
|
209
|
-
|
|
210
|
-
query: Query[Person] = Query.match(Person.name, Person.get_val(original_person))
|
|
211
|
-
|
|
212
|
-
found_people: List[PartialModel[Person]] = elastic_service.search_partial(query, source=[Person.name, Person.birth_date])
|
|
213
|
-
|
|
214
|
-
for found_person in found_people:
|
|
215
|
-
print(Person.name.get_val_unsafe(person))
|
|
216
|
-
# John
|
|
217
|
-
```
|
|
218
|
-
|
|
219
|
-
The above data model is built using the `ElasticModel` base class, which is a specialized subtype of `BaseModel` that captures property metadata specifically for Elasticsearch (e.g., the index name and field mappings). `Prop`s can be customized with `elastic_prop` to include field-level metadata like `mapping` (the `Prop` type has a metadata field for storing arbitrary key-value data for use cases like this).
|
|
220
|
-
|
|
221
|
-
`Query` is a representation of queries based on Pydoptic types. For instance, the match query is constructed by using a `Prop` value to specify the field to match with and providing a value corresponding to that prop's target type. The result, when querying `Person.name` is a `Query[Person]`. `Query[Person]` contains a reference to the `Person` class, which can be used to resolve the appropriate index name.
|
|
222
|
-
|
|
223
|
-
`ElasticService` provides an API for dealing with indices and documents using the Pydoptic-based `ElasticModel` and `Query` types. We first create our `Person` index by simply passing the class to `elastic_service.create_index`. The index name is generated by default from the class name (`person`) and the field mappings are constructed from the properties. We then index a document by passing it to `elastic_service.index`; since the instance contains a reference to the class, the index name can be resolved properly. Finally, we search for the original person by using our `Query[Person]`, which matches on the original person's name. When searching, however, we use a special variant `elastic_service.search_partial` which allows us to provide a `source` parameter, where we specify only two `Person` properties: `Person.name` and `Person.birth_date`. The result is a list of `PartialModel[Person]` instances containing only those fields.
|
|
224
|
-
|
|
225
|
-
You'll notice at no point in the above are we forced to pass any index or property names as strings. All the information required to resolve the indices and fields are contained in our model types.
|
|
226
|
-
|
|
227
|
-
### Reified optics
|
|
228
|
-
|
|
229
|
-
I hope the above has indicated clearly enough how pydoptic can be used to solve the problems indicated in the first section. At this point its worth saying a few words about the approach used.
|
|
230
|
-
|
|
231
|
-
Pydoptic is inspired by a concept from functional programming called "optics". In functional programming, optics are not just useful but pretty much necessary due to the relative difficulty of updating nested properties in immutable data structures. Since updating nested data is easier in imperative languages like Python, optics tend not to be used much. There a couple of optics libraries out there for Python, but they try to be functional in the full sense, which is to say they are designed to create immutable copies of dataclasses rather than mutate them.
|
|
232
|
-
|
|
233
|
-
Pydoptic's approach differs from most optics libraries in two ways. First, it takes the compositional properties of optic types from functional programming while giving them power to mutate objects. Second, it "reifies" the optics. Functional optics are typically encoded as *functions* that retrieve or update data and can be composed in various ways; in Pydoptic, they are encoded as *data*. They are *descriptions* of the things that *could be* accessed or updated. The actual functionality for doing the accessing/updating is secondary, and is implemented by *interpreting* the descriptions. This *reification* of the optics gives it a more general power than traditional optics.
|
|
234
|
-
|
|
235
|
-
This concept of reified optics comes from the Scala project [ZIO schema](https://zio.dev/zio-schema/), which provides (among other things) a similar mechanism for referencing properties via "accessors". Pydoptic provides a simplified version of this approach appropriate to Python and its more limited (though still quite powerful!) type system. The main innovation of Pydoptic (beyond bringing optics to mutable data) is that it models data "optics-first," unlike ZIO-schema, which (for good reasons connected with Scala and the JVM) *derives* optics *from* traditional data models (i.e., data classes).
|
|
236
|
-
|
|
237
14
|
## User guide
|
|
238
15
|
|
|
239
16
|
The examples below assume this import, unless a different one is shown:
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: pydoptic
|
|
3
|
-
Version: 0.0.
|
|
3
|
+
Version: 0.0.2.post1.dev1
|
|
4
4
|
Summary: Model data using reified optics
|
|
5
5
|
Author: John Hungerford
|
|
6
6
|
License-Expression: MIT
|
|
@@ -16,14 +16,12 @@ Provides-Extra: benchmark
|
|
|
16
16
|
Requires-Dist: pydantic~=2.11; extra == "benchmark"
|
|
17
17
|
Dynamic: license-file
|
|
18
18
|
|
|
19
|
-
#
|
|
19
|
+
# pydoptic
|
|
20
20
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
2. You frequently need to retrieve and manipulate deeply nested values
|
|
26
|
-
3. You need a type-safe way to refer to properties when querying remote data
|
|
21
|
+
The core data modeling library based on reified optics. See the
|
|
22
|
+
[repository README](https://github.com/johnhungerford/pydoptic) for what reified optics are and why
|
|
23
|
+
they're useful -- this document is the in-depth user guide: defining and constructing models, reading
|
|
24
|
+
and mutating data, chaining selects, discriminated subtypes, and integrating with external APIs.
|
|
27
25
|
|
|
28
26
|
## Installation
|
|
29
27
|
|
|
@@ -31,227 +29,6 @@ A good alternative to Pydantic when one or more of the following is true:
|
|
|
31
29
|
pip install pydoptic
|
|
32
30
|
```
|
|
33
31
|
|
|
34
|
-
This is the core data modeling library only. For the SQL query builder or the Elasticsearch
|
|
35
|
-
integration built on top of it, see [pydoptic-sql](https://pypi.org/project/pydoptic-sql/) and
|
|
36
|
-
[pydoptic-elastic](https://pypi.org/project/pydoptic-elastic/) -- both separate packages, published
|
|
37
|
-
from the same [repository](https://github.com/johnhungerford/pydoptic).
|
|
38
|
-
|
|
39
|
-
## Why?
|
|
40
|
-
|
|
41
|
-
In every programming language, data is modeled as the type of container used to store it in memory. In Python, this generally looks something like the following:
|
|
42
|
-
|
|
43
|
-
```python3
|
|
44
|
-
from pydantic import BaseModel
|
|
45
|
-
from datetime import date
|
|
46
|
-
from typing import List
|
|
47
|
-
|
|
48
|
-
class Address(BaseModel):
|
|
49
|
-
number: str
|
|
50
|
-
street: str
|
|
51
|
-
city: str
|
|
52
|
-
postal_code: str
|
|
53
|
-
state: str
|
|
54
|
-
|
|
55
|
-
class Organization(BaseModel):
|
|
56
|
-
name: str
|
|
57
|
-
address: Address
|
|
58
|
-
phone_number: str
|
|
59
|
-
owner: 'Person'
|
|
60
|
-
members: List['Person'] | None
|
|
61
|
-
|
|
62
|
-
class Person(BaseModel):
|
|
63
|
-
name: str
|
|
64
|
-
address: Address
|
|
65
|
-
birth_date: date
|
|
66
|
-
phone_number: str
|
|
67
|
-
organizations: List[Organization] | None
|
|
68
|
-
is_active: bool
|
|
69
|
-
```
|
|
70
|
-
|
|
71
|
-
While the above class definitions do provide a typed description of the data we expect to deal with, their utility when it comes to *doing* anything with the data is actually fairly limited. At most they can do the following:
|
|
72
|
-
|
|
73
|
-
1. Validate and then represent a complete instance of each data type, either from parameters or some serialized source (e.g., JSON)
|
|
74
|
-
2. Provide access to any property that can be validated with a type checker
|
|
75
|
-
3. Generate a valid serialized representation of an instance (e.g. JSON)
|
|
76
|
-
|
|
77
|
-
While this is certainly useful, consider all the other things we might want to do:
|
|
78
|
-
|
|
79
|
-
1. Consume and validate an *incomplete* instance of our data (e.g., a `Person`'s `name` and `birth_date`, but nothing else) without making the *complete* data model less precise (e.g., by making `Person.address` optional)
|
|
80
|
-
2. Access potentially missing values of an incomplete model
|
|
81
|
-
3. Query just the `name` and `birth_date` from a database. Nothing in the data model allows us to *specify the field itself* -- only the *value* of the field. This is an important distinction.
|
|
82
|
-
4. Update the `is_active` flag of every member of every organization a given person is a member of without having to check if `Person.organizations` or `Organization.members` are `None`.
|
|
83
|
-
|
|
84
|
-
Achieving these things with ordinary data models is inconvenient at best. Dealing with incomplete data requires defining new types for each use case or abandoning type safety altogether. While data models can read and write (complete) records to a datastore, they provide no mechanism for helping us specify the parts of the data model that are of interest to us, like when we want to specify which fields to retrieve or define a query setting some constraint on a particular field. In general we have to specify field names as strings, losing any pre-runtime validation. Finally, traditional data models provide no abstractions for manipulating nested data, nor do they provide any mechanism that could be used to create such abstractions.
|
|
85
|
-
|
|
86
|
-
Pydoptic provides an alternative way to model data in Python -- and indeed in any language -- that allows you to accomplish all these and things much more easily and safely.
|
|
87
|
-
|
|
88
|
-
## How?
|
|
89
|
-
|
|
90
|
-
Whereas traditional data models represent data as *containers* of properties, Pydoptic models data as collections of *references* to properties. In a traditional data type, a "property" is a value contained by an in-memory instance of the type in question. In Pydoptic, a "property" is a *description* of a value belonging to a type. This description can be used to do the usual things -- storing a value in an object or retrieving it from an object -- but it can also do a much more. For instance, it can be used to store a value in a *remote* data source, or query it and retrieve it from that data source, or... pretty much anything else you can think of that might need to be done with it!
|
|
91
|
-
|
|
92
|
-
Here is what a Pydoptic version of the data model above looks like:
|
|
93
|
-
|
|
94
|
-
```python3
|
|
95
|
-
from pydoptic import BaseModel, PartialModel, Prop, PropOpt, PropOptArr
|
|
96
|
-
from datetime import date
|
|
97
|
-
|
|
98
|
-
class Address(BaseModel):
|
|
99
|
-
number: Prop['Address', str]
|
|
100
|
-
street: Prop['Address', str]
|
|
101
|
-
city: Prop['Address', str]
|
|
102
|
-
postal_code: Prop['Address', str]
|
|
103
|
-
state: Prop['Address', str]
|
|
104
|
-
|
|
105
|
-
class Organization(BaseModel):
|
|
106
|
-
name: Prop['Organization', str]
|
|
107
|
-
address: Prop['Organization', Address]
|
|
108
|
-
phone_number: Prop['Organization', str]
|
|
109
|
-
owner: PropOpt['Organization', 'Person']
|
|
110
|
-
members: PropOptArr['Organization', 'Person']
|
|
111
|
-
|
|
112
|
-
class Person(BaseModel):
|
|
113
|
-
name: Prop['Person', str]
|
|
114
|
-
address: Prop['Person', Address]
|
|
115
|
-
birth_date: Prop['Person', date]
|
|
116
|
-
phone_number: Prop['Person', str]
|
|
117
|
-
organizations: PropOptArr['Person', Organization]
|
|
118
|
-
is_active: Prop['Person', bool]
|
|
119
|
-
```
|
|
120
|
-
|
|
121
|
-
The type signatures of the properties include some more boilerplate, to be sure. Most notably, they all contain references to the model class that they belong to. This may seem redundant, but it's a crucial feature that makes them as powerful as the are: because they contain references back to the classes they belong to, they can be used entirely independently of those classes. Each property -- which is only a *class* attribute -- is initialized with a value containing references to both its "origin" type (the model class) and its "target" type (the second type parameter on the right) along with its attribute name and flags indicating whether it's optional (for `PropOpt` properties), array (`PropArr`), or both (`PropOptArr`).
|
|
122
|
-
|
|
123
|
-
By constructing our properties as comprehensive metadata *about* values and their relationship with the model they belong to, rather than simply the values themselves, we provide ourselves with a much more flexible and powerful tool. Let's see what we can do with them.
|
|
124
|
-
|
|
125
|
-
### Basics
|
|
126
|
-
|
|
127
|
-
Let's start with the basics. These "props" would not be much use to us if we could not actually construct model instances. It turns out we can do this in the usual way:
|
|
128
|
-
|
|
129
|
-
```python3
|
|
130
|
-
person = Person(name="John", address=Address(...), birth_date=..., phone_number=..., is_active=True)
|
|
131
|
-
|
|
132
|
-
print(person.name)
|
|
133
|
-
# John
|
|
134
|
-
```
|
|
135
|
-
|
|
136
|
-
On initialization, the model class constructs the `Prop` class attributes based on the type hints and keeps track of them internally. It then uses the known properties to validate keyword arguments that are provided when constructing a class instance. The above succeeds even though `organizations` is missing because organizations is a `PropOptArr`, which is optional. If we left out `name`, however, it would raise a `ValueError`.
|
|
137
|
-
|
|
138
|
-
Note that your type checker will complain that `person.name` is a `Prop` instead of a `str`. While Pydoptic does assign properties as model attributes (the `Prop` types should be defined only on the class), your type checker will not know this. The "Pydoptic" way to retrieve properties is not to access the attribute directly, but use the property itself!
|
|
139
|
-
|
|
140
|
-
```python3
|
|
141
|
-
person_name = Person.name.get_val(person)
|
|
142
|
-
|
|
143
|
-
print(person_name)
|
|
144
|
-
# John
|
|
145
|
-
```
|
|
146
|
-
|
|
147
|
-
Your type checker will recognize `person_name` as having a type `str`. While this is a fairly verbose way of getting a simple value, its utility will become clearer when you find yourself frequently accessing nested data. For instance, say you want to flip the `is_active` status of every members of every organization a given person is connected with. Ordinarily this would require a fairly elaborate combination of `for` loops and and `if` statements:
|
|
148
|
-
|
|
149
|
-
```python3
|
|
150
|
-
if person.organizations is not None:
|
|
151
|
-
for organization in organizations:
|
|
152
|
-
if organization.members is not None:
|
|
153
|
-
for member in members:
|
|
154
|
-
member.is_active = !member.is_active
|
|
155
|
-
```
|
|
156
|
-
|
|
157
|
-
In this case, Pydoptic's model is substantially *less* verbose. We can reproduce all of the above logic by *composing* our `Prop`s to select the desired path to `is_active` and then use the `update` method to change every value:
|
|
158
|
-
|
|
159
|
-
```python3
|
|
160
|
-
select_related_statuses = Person.organizations(Organization.members)(Person.is_active)
|
|
161
|
-
|
|
162
|
-
statuses = select_related_statuses.get(person).as_list
|
|
163
|
-
print(statuses)
|
|
164
|
-
# [True, False, False, True, ...]
|
|
165
|
-
|
|
166
|
-
# All the updating is done here:
|
|
167
|
-
select_related_statuses.update(person, lambda status: !status)
|
|
168
|
-
|
|
169
|
-
statuses = select_related_statuses.get(person).as_list
|
|
170
|
-
print(statuses)
|
|
171
|
-
# [False, True, True, False, ...]
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
As your type checker should indicate, `select_related_statuses` is a `Select[Person, bool]` which is composed from the `Prop`s `Person.organizations`, `Organization.members`, and `Person.is_active`. This particular chain of props can compose because each prop's *target* type is the same as the *origin* type of the prop chained to it. When this is not the case (meaning the chaining is *invalid*), the type checker will complain. The resulting `Select` can then be used to retrieve and update all `Person.is_active` properties (the final `Prop` in the chain) nested within the original `person` via the path `organizations` -> `members` -> `is_active`. Since `Person.organizations` and `Organization.members` are optional properties, ordinarily retrieving and updating these values would require checking for `None` in multiple places.
|
|
175
|
-
|
|
176
|
-
### Incomplete data
|
|
177
|
-
|
|
178
|
-
By separating our property types from the actual container, handling incomplete data becomes simpler without losing type precision. The same `Prop`s we use to set and retrieve data a model like `Person` can be used to do the same in incomplete models as well.
|
|
179
|
-
|
|
180
|
-
```python3
|
|
181
|
-
partial_person: PartialModel[Person] = Person.partial(name="John", birth_date=...)
|
|
182
|
-
|
|
183
|
-
print(Person.name.get_val_unsafe(partial_person))
|
|
184
|
-
# John
|
|
185
|
-
|
|
186
|
-
person_dict = {'name': 'John'}
|
|
187
|
-
print(Person.name.get_val_unsafe(person_dict))
|
|
188
|
-
# John
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
As you can see in the example above, there are two ways to represent incomplete data. `PartialModel` is a model type that can have missing properties and extraneous properties, but any property whose name corresponds to a property on the full model must be valid. Hence `Person.partial(name=23)` would fail because `23` is not a string. This is useful when you are consuming partial data from a data source but still want to validate it. If you don't want to validate the data at all, you can also use a `dict`. If you know your data source is producing valid data, it probably makes sense to just keep your data in an untyped `Dict[str, Any]`.
|
|
192
|
-
|
|
193
|
-
All of the data manipulation methods used on full models have versions that can be used on `PartialModel`s and `dict`s. Methods with the suffix `_unsafe` will work exactly like the regular method, but will raise a `ValueError` when required data is missing (or invalid in certain ways); methods with the suffix `_safe` will return `None` when data is invalid or (for mutating methods) fail silently.
|
|
194
|
-
|
|
195
|
-
### Data integration
|
|
196
|
-
|
|
197
|
-
Since `Prop`s contain all the information about the properties they reference, they can easily be used for interacting with *remote* instances of the same data. This repository includes an example Elasticsearch integration as a reference. Here is what it looks like to use the pydoptic-based Elasticsearch API:
|
|
198
|
-
|
|
199
|
-
```python3
|
|
200
|
-
from pydoptic import Prop, PropOptArr, PartialModel
|
|
201
|
-
from pydoptic_elastic import ElasticModel, Query, ElasticService, elastic_prop, ESMapping
|
|
202
|
-
from datetime import date
|
|
203
|
-
from typing import List
|
|
204
|
-
from elasticsearch import Elasticsearch
|
|
205
|
-
|
|
206
|
-
class Address(ElasticModel):
|
|
207
|
-
...
|
|
208
|
-
|
|
209
|
-
class Organization(ElasticModel):
|
|
210
|
-
...
|
|
211
|
-
|
|
212
|
-
class Person(ElasticModel):
|
|
213
|
-
name: Prop['Person', str]
|
|
214
|
-
address: Prop['Person', Address]
|
|
215
|
-
birth_date: Prop['Person', date]
|
|
216
|
-
phone_number: Prop['Person', str] = elastic_prop(mapping=ESMapping.keyword)
|
|
217
|
-
organizations: PropOptArr['Person', Organization]
|
|
218
|
-
is_active: Prop['Person', bool]
|
|
219
|
-
|
|
220
|
-
elastic_service = ElasticService(Elasticsearch('http://localhost:9200'))
|
|
221
|
-
|
|
222
|
-
elastic_service.create_index(Person)
|
|
223
|
-
|
|
224
|
-
original_person = Person(name='John', ...)
|
|
225
|
-
|
|
226
|
-
elastic_service.index(original_person)
|
|
227
|
-
|
|
228
|
-
query: Query[Person] = Query.match(Person.name, Person.get_val(original_person))
|
|
229
|
-
|
|
230
|
-
found_people: List[PartialModel[Person]] = elastic_service.search_partial(query, source=[Person.name, Person.birth_date])
|
|
231
|
-
|
|
232
|
-
for found_person in found_people:
|
|
233
|
-
print(Person.name.get_val_unsafe(person))
|
|
234
|
-
# John
|
|
235
|
-
```
|
|
236
|
-
|
|
237
|
-
The above data model is built using the `ElasticModel` base class, which is a specialized subtype of `BaseModel` that captures property metadata specifically for Elasticsearch (e.g., the index name and field mappings). `Prop`s can be customized with `elastic_prop` to include field-level metadata like `mapping` (the `Prop` type has a metadata field for storing arbitrary key-value data for use cases like this).
|
|
238
|
-
|
|
239
|
-
`Query` is a representation of queries based on Pydoptic types. For instance, the match query is constructed by using a `Prop` value to specify the field to match with and providing a value corresponding to that prop's target type. The result, when querying `Person.name` is a `Query[Person]`. `Query[Person]` contains a reference to the `Person` class, which can be used to resolve the appropriate index name.
|
|
240
|
-
|
|
241
|
-
`ElasticService` provides an API for dealing with indices and documents using the Pydoptic-based `ElasticModel` and `Query` types. We first create our `Person` index by simply passing the class to `elastic_service.create_index`. The index name is generated by default from the class name (`person`) and the field mappings are constructed from the properties. We then index a document by passing it to `elastic_service.index`; since the instance contains a reference to the class, the index name can be resolved properly. Finally, we search for the original person by using our `Query[Person]`, which matches on the original person's name. When searching, however, we use a special variant `elastic_service.search_partial` which allows us to provide a `source` parameter, where we specify only two `Person` properties: `Person.name` and `Person.birth_date`. The result is a list of `PartialModel[Person]` instances containing only those fields.
|
|
242
|
-
|
|
243
|
-
You'll notice at no point in the above are we forced to pass any index or property names as strings. All the information required to resolve the indices and fields are contained in our model types.
|
|
244
|
-
|
|
245
|
-
### Reified optics
|
|
246
|
-
|
|
247
|
-
I hope the above has indicated clearly enough how pydoptic can be used to solve the problems indicated in the first section. At this point its worth saying a few words about the approach used.
|
|
248
|
-
|
|
249
|
-
Pydoptic is inspired by a concept from functional programming called "optics". In functional programming, optics are not just useful but pretty much necessary due to the relative difficulty of updating nested properties in immutable data structures. Since updating nested data is easier in imperative languages like Python, optics tend not to be used much. There a couple of optics libraries out there for Python, but they try to be functional in the full sense, which is to say they are designed to create immutable copies of dataclasses rather than mutate them.
|
|
250
|
-
|
|
251
|
-
Pydoptic's approach differs from most optics libraries in two ways. First, it takes the compositional properties of optic types from functional programming while giving them power to mutate objects. Second, it "reifies" the optics. Functional optics are typically encoded as *functions* that retrieve or update data and can be composed in various ways; in Pydoptic, they are encoded as *data*. They are *descriptions* of the things that *could be* accessed or updated. The actual functionality for doing the accessing/updating is secondary, and is implemented by *interpreting* the descriptions. This *reification* of the optics gives it a more general power than traditional optics.
|
|
252
|
-
|
|
253
|
-
This concept of reified optics comes from the Scala project [ZIO schema](https://zio.dev/zio-schema/), which provides (among other things) a similar mechanism for referencing properties via "accessors". Pydoptic provides a simplified version of this approach appropriate to Python and its more limited (though still quite powerful!) type system. The main innovation of Pydoptic (beyond bringing optics to mutable data) is that it models data "optics-first," unlike ZIO-schema, which (for good reasons connected with Scala and the JVM) *derives* optics *from* traditional data models (i.e., data classes).
|
|
254
|
-
|
|
255
32
|
## User guide
|
|
256
33
|
|
|
257
34
|
The examples below assume this import, unless a different one is shown:
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
{pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/dependency_links.txt
RENAMED
|
File without changes
|
|
File without changes
|
{pydoptic-0.0.1.post1.dev2 → pydoptic-0.0.2.post1.dev1}/src/pydoptic.egg-info/scm_file_list.json
RENAMED
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|