checkpointer 2.14.11__tar.gz → 2.14.12__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checkpointer-2.14.12/PKG-INFO +385 -0
- checkpointer-2.14.12/README.md +368 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/storages/pickle_storage.py +6 -8
- {checkpointer-2.14.11 → checkpointer-2.14.12}/pyproject.toml +1 -1
- {checkpointer-2.14.11 → checkpointer-2.14.12}/uv.lock +1 -1
- checkpointer-2.14.11/PKG-INFO +0 -262
- checkpointer-2.14.11/README.md +0 -245
- {checkpointer-2.14.11 → checkpointer-2.14.12}/.gitignore +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/.python-version +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/ATTRIBUTION.md +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/LICENSE +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/__init__.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/checkpoint.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/fn_ident.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/fn_string.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/import_mappings.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/object_hash.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/print_checkpoint.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/storages/__init__.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/storages/memory_storage.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/storages/storage.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/types.py +0 -0
- {checkpointer-2.14.11 → checkpointer-2.14.12}/checkpointer/utils.py +0 -0
|
@@ -0,0 +1,385 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: checkpointer
|
|
3
|
+
Version: 2.14.12
|
|
4
|
+
Summary: checkpointer adds code-aware caching to Python functions, maintaining correctness and speeding up execution as your code changes.
|
|
5
|
+
Project-URL: Repository, https://github.com/Reddan/checkpointer.git
|
|
6
|
+
Author: Hampus Hallman
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
License-File: ATTRIBUTION.md
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: async,cache,caching,data analysis,data processing,fast,hashing,invalidation,memoization,optimization,performance,workflow
|
|
11
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
15
|
+
Requires-Python: >=3.11
|
|
16
|
+
Description-Content-Type: text/markdown
|
|
17
|
+
|
|
18
|
+
# checkpointer · [](https://github.com/Reddan/checkpointer/blob/master/LICENSE) [](https://pypi.org/project/checkpointer/) [](https://pypi.org/project/checkpointer/)
|
|
19
|
+
|
|
20
|
+
`checkpointer` is a Python library for memoizing function results with **code-aware cache invalidation**. Decorate a function with `@checkpoint` and its return values are cached to disk (or memory). When you call it again with the same arguments, the cached result is returned instead of recomputing it.
|
|
21
|
+
|
|
22
|
+
What makes it different from ordinary memoization is that the cache invalidates itself automatically when your **code** changes - not just when arguments change. Edit a function's logic, or the logic of anything it depends on, and the stale cache is discarded on the next run. You get the speed of caching without the classic footgun of serving results from code that no longer exists.
|
|
23
|
+
|
|
24
|
+
It works with sync and async functions, methods, and recursion, handles complex objects and large **NumPy** / **PyTorch** arrays, and lets you fine-tune exactly what counts toward a cache key.
|
|
25
|
+
|
|
26
|
+
## 📦 Installation
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
pip install checkpointer
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Requires Python 3.11+. No mandatory dependencies. NumPy, PyTorch, and Polars are supported automatically if they happen to be installed.
|
|
33
|
+
|
|
34
|
+
## 🚀 Quick Start
|
|
35
|
+
|
|
36
|
+
```python
|
|
37
|
+
from checkpointer import checkpoint
|
|
38
|
+
|
|
39
|
+
@checkpoint
|
|
40
|
+
def load_dataset(path: str) -> pl.DataFrame:
|
|
41
|
+
print("Reading and parsing...")
|
|
42
|
+
return pl.read_csv(path).filter(pl.col("price") > 0)
|
|
43
|
+
|
|
44
|
+
df = load_dataset("sales.csv") # Reads, parses, and caches the DataFrame
|
|
45
|
+
df = load_dataset("sales.csv") # Skips the work - loaded from cache
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
The win shows up in everyday iteration: a script or notebook that reloads a multi-second dataset on every run reads it once and reuses the result on subsequent runs. And because the cache is *code-aware*, the moment you change `load_dataset` (say, tighten the filter), the stale result is dropped and the file is re-parsed - no manual cache-busting.
|
|
49
|
+
|
|
50
|
+
By default, results are pickled to `~/.cache/checkpoints`, so the cache survives across processes and restarts. (Polars `DataFrame`s are stored as Parquet automatically.)
|
|
51
|
+
|
|
52
|
+
## 🧠 How It Works
|
|
53
|
+
|
|
54
|
+
Every cached call is identified by two hashes:
|
|
55
|
+
|
|
56
|
+
- **Function identity hash** - computed once per function (on first use). It captures the function's *source code* and the source of every user-defined function, method, and class it depends on, recursively. Change any of that logic and the hash changes, invalidating all cached results for that function. Cosmetic edits (comments, whitespace, formatting, type annotations) are deliberately ignored.
|
|
57
|
+
- **Call hash** - computed on every call from the actual arguments (and, optionally, captured global variables). Different arguments produce different call hashes.
|
|
58
|
+
|
|
59
|
+
When you call a decorated function, `checkpointer` combines these into a lookup key. If a valid cached result exists, it's returned immediately; otherwise the function runs, the result is stored, and then returned.
|
|
60
|
+
|
|
61
|
+
Because dependency tracking is automatic, you rarely need to bump a version number by hand - editing the code *is* the version bump.
|
|
62
|
+
|
|
63
|
+
### What counts as a dependency
|
|
64
|
+
|
|
65
|
+
The identity hash follows your function into the code it actually uses. `checkpointer` discovers dependencies by:
|
|
66
|
+
|
|
67
|
+
- **Inspecting the global scope** - functions, methods, and classes the function references are pulled in (recursively, including their dependencies).
|
|
68
|
+
- **Inferring from type annotations** - classes named in argument annotations are treated as dependencies, so changes to their methods invalidate the cache too.
|
|
69
|
+
- **Analyzing constructions and calls** - objects built and methods invoked inside the function are traced back to the classes and methods they come from.
|
|
70
|
+
|
|
71
|
+
A few things deliberately **don't** invalidate:
|
|
72
|
+
|
|
73
|
+
- Cosmetic edits - comments, whitespace, formatting, and parameter type annotations.
|
|
74
|
+
- Changes elsewhere in the module that the function doesn't touch.
|
|
75
|
+
- Changing a parameter's default value, unless it changes the actual arguments a call resolves to.
|
|
76
|
+
|
|
77
|
+
## 💡 Examples
|
|
78
|
+
|
|
79
|
+
### Async functions
|
|
80
|
+
|
|
81
|
+
Works with any async runtime - the awaited value is what gets cached, so repeated calls skip the network entirely.
|
|
82
|
+
|
|
83
|
+
```python
|
|
84
|
+
@checkpoint
|
|
85
|
+
async def fetch_profile(user_id: int) -> dict:
|
|
86
|
+
async with httpx.AsyncClient() as client:
|
|
87
|
+
resp = await client.get(f"https://api.example.com/users/{user_id}")
|
|
88
|
+
return resp.json()
|
|
89
|
+
|
|
90
|
+
profile = await fetch_profile(42) # Hits the API
|
|
91
|
+
profile = await fetch_profile(42) # Instant - from cache
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
### Methods
|
|
95
|
+
|
|
96
|
+
Decorate methods directly. The instance is hashed as part of the call, so results are keyed to it - two embedders with different models cache separately, with no collisions.
|
|
97
|
+
|
|
98
|
+
Two things make this example efficient: the method returns the NumPy array **as-is** (don't `.tolist()` it - `checkpointer` pickles arrays compactly and far faster than a Python list), and the class defines `__objecthash__` so hashing an instance is instant instead of crawling the whole loaded model. See [Custom Instance Hashing](#custom-instance-hashing-with-__objecthash__) for the details.
|
|
99
|
+
|
|
100
|
+
```python
|
|
101
|
+
import numpy as np
|
|
102
|
+
from sentence_transformers import SentenceTransformer
|
|
103
|
+
|
|
104
|
+
class Embedder:
|
|
105
|
+
def __init__(self, model_name: str):
|
|
106
|
+
self.model_name = model_name
|
|
107
|
+
self.model = SentenceTransformer(model_name) # Loaded once, reused
|
|
108
|
+
|
|
109
|
+
def __objecthash__(self):
|
|
110
|
+
return self.model_name # Fast, stable identity - skips hashing the model
|
|
111
|
+
|
|
112
|
+
@checkpoint
|
|
113
|
+
def embed(self, text: str) -> np.ndarray:
|
|
114
|
+
return self.model.encode(text) # Cached as a NumPy array
|
|
115
|
+
|
|
116
|
+
fast = Embedder("all-MiniLM-L6-v2")
|
|
117
|
+
fast.embed("hello world") # Computed and cached for this model
|
|
118
|
+
fast.embed("hello world") # From cache - the model isn't even consulted
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
### Force recomputation
|
|
122
|
+
|
|
123
|
+
`.rerun(...)` runs the function and overwrites the cache - useful when an upstream data source changed but your code didn't.
|
|
124
|
+
|
|
125
|
+
```python
|
|
126
|
+
df = load_dataset("sales.csv") # Cached
|
|
127
|
+
df = load_dataset.rerun("sales.csv") # Recomputes and overwrites the cache
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
### Expiry / TTL
|
|
131
|
+
|
|
132
|
+
Expire results by age with a `timedelta`, or by a custom rule with a callable that receives the store timestamp and returns `True` when stale.
|
|
133
|
+
|
|
134
|
+
```python
|
|
135
|
+
from datetime import datetime, timedelta
|
|
136
|
+
|
|
137
|
+
# Re-fetch a volatile rate at most once every 15 minutes
|
|
138
|
+
@checkpoint(expiry=timedelta(minutes=15))
|
|
139
|
+
def get_exchange_rate(base: str, quote: str) -> float:
|
|
140
|
+
return httpx.get(f"https://api.example.com/rate/{base}/{quote}").json()["rate"]
|
|
141
|
+
|
|
142
|
+
# Invalidate anything cached before today's UTC midnight
|
|
143
|
+
@checkpoint(expiry=lambda stored_at: stored_at.date() < datetime.utcnow().date())
|
|
144
|
+
def daily_report(team: str) -> dict: ...
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
### Layered / multi-backend caching
|
|
148
|
+
|
|
149
|
+
Stack decorators to combine backends - e.g. a fast in-memory layer in front of a persistent disk layer - without losing cache consistency. Great for a lookup hit many times per run that's also worth keeping across runs.
|
|
150
|
+
|
|
151
|
+
```python
|
|
152
|
+
@checkpoint(storage="memory") # Hot path, in-process
|
|
153
|
+
@checkpoint(storage="pickle") # Persistent, on disk
|
|
154
|
+
def geocode(address: str) -> tuple[float, float]:
|
|
155
|
+
resp = httpx.get("https://api.example.com/geocode", params={"q": address})
|
|
156
|
+
return tuple(resp.json()["latlng"])
|
|
157
|
+
|
|
158
|
+
geocode("1600 Amphitheatre Pkwy") # API call, written to both layers
|
|
159
|
+
geocode("1600 Amphitheatre Pkwy") # From memory
|
|
160
|
+
geocode.fn.get("1600 Amphitheatre Pkwy") # From the pickle layer underneath
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
### Toggle caching on/off
|
|
164
|
+
|
|
165
|
+
Flip caching with `when` - keep the persisted cache while iterating locally, but run clean in production (or in tests).
|
|
166
|
+
|
|
167
|
+
```python
|
|
168
|
+
import os
|
|
169
|
+
|
|
170
|
+
IS_DEV = os.environ.get("ENV") == "dev"
|
|
171
|
+
|
|
172
|
+
@checkpoint(when=IS_DEV) # Caches while developing; runs straight through otherwise
|
|
173
|
+
def build_features(df: pl.DataFrame) -> pl.DataFrame:
|
|
174
|
+
return df.with_columns(...) # expensive feature engineering
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
### Recursion
|
|
178
|
+
|
|
179
|
+
Inside a recursive function, call `.fn(...)` to invoke the original, undecorated function. This caches the top-level result without writing a separate checkpoint for every intermediate step - handy when the recursion fans out over expensive calls.
|
|
180
|
+
|
|
181
|
+
```python
|
|
182
|
+
@checkpoint
|
|
183
|
+
def resolve_deps(package: str) -> set[str]:
|
|
184
|
+
deps = fetch_dependencies(package) # e.g. a registry API call
|
|
185
|
+
return deps | {sub for dep in deps for sub in resolve_deps.fn(dep)}
|
|
186
|
+
|
|
187
|
+
resolve_deps("flask") # Caches the fully-resolved dependency set
|
|
188
|
+
resolve_deps.get("flask") # Reads it back; transitive deps weren't cached individually
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
## Customizing How Arguments Are Hashed
|
|
192
|
+
|
|
193
|
+
Control what an argument contributes to the call hash - without changing the value the function actually receives. Useful for normalization (better hit rates) or for hashing something cheaper/more meaningful than the raw object.
|
|
194
|
+
|
|
195
|
+
- **`Annotated[T, HashBy[fn]]`** - hash `fn(arg)` instead of `arg`.
|
|
196
|
+
- **`NoHash[T]`** - exclude the argument from the hash entirely.
|
|
197
|
+
|
|
198
|
+
```python
|
|
199
|
+
from typing import Annotated
|
|
200
|
+
from pathlib import Path
|
|
201
|
+
import logging
|
|
202
|
+
from checkpointer import checkpoint, HashBy, NoHash
|
|
203
|
+
|
|
204
|
+
def file_bytes(path: Path) -> bytes:
|
|
205
|
+
return path.read_bytes()
|
|
206
|
+
|
|
207
|
+
@checkpoint
|
|
208
|
+
def process(
|
|
209
|
+
numbers: Annotated[list[int], HashBy[sorted]], # Order-insensitive
|
|
210
|
+
data_file: Annotated[Path, HashBy[file_bytes]], # Hash by file contents, not path
|
|
211
|
+
log: NoHash[logging.Logger], # Ignored entirely
|
|
212
|
+
):
|
|
213
|
+
...
|
|
214
|
+
```
|
|
215
|
+
|
|
216
|
+
Here `[3, 1, 2]` and `[1, 2, 3]` hit the same cache entry, the cache tracks the file's *contents* rather than its name, and swapping loggers never invalidates anything.
|
|
217
|
+
|
|
218
|
+
## Custom Instance Hashing with `__objecthash__`
|
|
219
|
+
|
|
220
|
+
Any class can define `__objecthash__` to control how its instances are hashed. When `checkpointer` encounters an instance, it hashes the return value of `__objecthash__()` instead of inspecting the object's internals.
|
|
221
|
+
|
|
222
|
+
```python
|
|
223
|
+
class Model:
|
|
224
|
+
def __init__(self, id: str, weights: list[float]):
|
|
225
|
+
self.id = id
|
|
226
|
+
self.weights = weights
|
|
227
|
+
|
|
228
|
+
def __objecthash__(self):
|
|
229
|
+
return self.id # Identity depends only on `id`
|
|
230
|
+
```
|
|
231
|
+
|
|
232
|
+
The return value can be anything `checkpointer` knows how to hash - a string, tuple, dict, etc. Once defined, it applies everywhere the class appears: as an argument, a captured variable, or nested inside another value - no per-call-site annotation needed.
|
|
233
|
+
|
|
234
|
+
## Capturing Global Variables
|
|
235
|
+
|
|
236
|
+
Sometimes a function's result depends on a module-level global, not just its arguments. `checkpointer` can fold such **captured globals** into the call hash so the cache invalidates when they change.
|
|
237
|
+
|
|
238
|
+
Enable it broadly with `capture=True` (captures every referenced global except those marked `NoHash`), or opt in per-variable with annotations:
|
|
239
|
+
|
|
240
|
+
- **`CaptureMe[T]`** - hashed on *every* call; changes invalidate immediately.
|
|
241
|
+
- **`CaptureMeOnce[T]`** - hashed *once per Python session*; cheaper, for expensive immutable globals.
|
|
242
|
+
|
|
243
|
+
Both combine with `HashBy` to customize hashing.
|
|
244
|
+
|
|
245
|
+
```python
|
|
246
|
+
from typing import Annotated
|
|
247
|
+
from pathlib import Path
|
|
248
|
+
from checkpointer import checkpoint, CaptureMe, CaptureMeOnce, HashBy
|
|
249
|
+
|
|
250
|
+
def file_bytes(path: Path) -> bytes:
|
|
251
|
+
return path.read_bytes()
|
|
252
|
+
|
|
253
|
+
config_file: CaptureMe[Annotated[Path, HashBy[file_bytes]]] = Path("config.yaml")
|
|
254
|
+
session_seed: CaptureMeOnce[int] = 42
|
|
255
|
+
|
|
256
|
+
@checkpoint
|
|
257
|
+
def run():
|
|
258
|
+
# Re-hashes `config_file` (by contents) every call;
|
|
259
|
+
# hashes `session_seed` once per session.
|
|
260
|
+
...
|
|
261
|
+
```
|
|
262
|
+
|
|
263
|
+
## Custom Storage Backends
|
|
264
|
+
|
|
265
|
+
Beyond the built-in `"pickle"` and `"memory"` backends, you can implement your own - e.g. to cache in Redis, S3, or a database. Subclass `Storage`, implement a handful of methods, and pass the class as `storage`. Calls are identified by `call_hash`; use `self.fn_id()` to namespace entries by function identity (name + version hash).
|
|
266
|
+
|
|
267
|
+
```python
|
|
268
|
+
from checkpointer import checkpoint, Storage
|
|
269
|
+
|
|
270
|
+
class RedisStorage(Storage):
|
|
271
|
+
def store(self, call_hash, data):
|
|
272
|
+
redis.set(self._key(call_hash), pickle.dumps(data))
|
|
273
|
+
return data # must return data
|
|
274
|
+
def load(self, call_hash):
|
|
275
|
+
return pickle.loads(redis.get(self._key(call_hash)))
|
|
276
|
+
def exists(self, call_hash):
|
|
277
|
+
return bool(redis.exists(self._key(call_hash)))
|
|
278
|
+
# ...plus delete() and checkpoint_date()
|
|
279
|
+
|
|
280
|
+
@checkpoint(storage=RedisStorage)
|
|
281
|
+
def cached(x: int):
|
|
282
|
+
return x ** 2
|
|
283
|
+
```
|
|
284
|
+
|
|
285
|
+
See [the `Storage` interface](#custom-storage-interface) in the API reference for the complete set of methods.
|
|
286
|
+
|
|
287
|
+
---
|
|
288
|
+
|
|
289
|
+
# 📚 API Reference
|
|
290
|
+
|
|
291
|
+
## `@checkpoint`
|
|
292
|
+
|
|
293
|
+
The default decorator. Also available as a configurable factory - call it with options to get a new, reusable checkpointer:
|
|
294
|
+
|
|
295
|
+
```python
|
|
296
|
+
@checkpoint # use defaults
|
|
297
|
+
@checkpoint(storage="memory") # override options
|
|
298
|
+
dev = checkpoint(when=IS_DEV) # reusable preset
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
### Options
|
|
302
|
+
|
|
303
|
+
| Option | Type | Default | Description |
|
|
304
|
+
| --- | --- | --- | --- |
|
|
305
|
+
| `storage` | `"pickle"` \| `"memory"` \| `type[Storage]` | `"pickle"` | Backend. `"pickle"` is persistent on disk; `"memory"` lives in-process; or pass a custom `Storage` subclass. |
|
|
306
|
+
| `directory` | `str` \| `Path` \| `None` | `~/.cache/checkpoints` | Root directory for the `"pickle"` backend. |
|
|
307
|
+
| `capture` | `bool` | `False` | If `True`, include all referenced globals in call hashes (except those marked `NoHash`). |
|
|
308
|
+
| `expiry` | `timedelta` \| `Callable[[datetime], bool]` \| `None` | `None` | Treat a cached result as stale. A `timedelta` expires by age; a callable receives the store timestamp and returns `True` when expired. |
|
|
309
|
+
| `fn_hash_from` | `Any` | `None` | Override the computed function-identity hash with any hashable value (a version string, config id, etc.). Set this and source-code changes no longer auto-invalidate - *you* control the version. |
|
|
310
|
+
| `when` | `bool` | `True` | Master on/off switch. When `False`, calls run straight through with no caching. |
|
|
311
|
+
| `verbosity` | `0` \| `1` \| `2` | `1` | `0`: silent. `1`: log on compute/store. `2`: also log on cache hits. |
|
|
312
|
+
|
|
313
|
+
## `CachedFunction` methods
|
|
314
|
+
|
|
315
|
+
A decorated function becomes a `CachedFunction`. Calling it normally caches or loads; the following give finer control. (`*args, **kw` below are always the function's own arguments.)
|
|
316
|
+
|
|
317
|
+
| Member | Description |
|
|
318
|
+
| --- | --- |
|
|
319
|
+
| `fn(*args, **kw)` | The original, undecorated function (a property). Bypasses the cache - use it in recursion. |
|
|
320
|
+
| `rerun(*args, **kw)` | Force execution and overwrite any cached result. |
|
|
321
|
+
| `cached(*args, **kw)` | Like calling normally, but ignores `when=False` (always uses the cache). |
|
|
322
|
+
| `get(*args, **kw)` | Return the cached result without computing. Raises `CheckpointError` if absent. |
|
|
323
|
+
| `get_or(default, *args, **kw)` | Like `get`, but returns `default` instead of raising. |
|
|
324
|
+
| `set(value, *args, **kw)` | Manually store `value` as the result for these arguments. Use this for a **sync** function, whose return value is stored directly. |
|
|
325
|
+
| `set_awaitable(value, *args, **kw)` | The `set` for an **async** function, whose resolved value is stored wrapped so that loading it yields an awaitable (matching the original signature) - the wrapping is handled for you. |
|
|
326
|
+
| `exists(*args, **kw)` | `True` if a cached entry exists for these arguments. |
|
|
327
|
+
| `delete(*args, **kw)` | Remove the cached entry for these arguments. |
|
|
328
|
+
| `get_call_hash(*args, **kw)` | The call hash these arguments produce. |
|
|
329
|
+
| `is_expired(call_hash)` | `True` if no entry exists or it has expired per `expiry`. |
|
|
330
|
+
| `reinit(recursive=True)` | Recompute the function-identity hash and re-capture `CaptureMeOnce` globals within the current session. |
|
|
331
|
+
| `cleanup(invalidated=True, expired=True)` | Delete checkpoints from outdated function versions and/or expired entries. |
|
|
332
|
+
| `ident` | The `FunctionIdent` - exposes `fn_hash`, dependencies, and capturables. |
|
|
333
|
+
| `storage` | The bound `Storage` instance. |
|
|
334
|
+
|
|
335
|
+
## Annotations & types
|
|
336
|
+
|
|
337
|
+
Importable from `checkpointer`:
|
|
338
|
+
|
|
339
|
+
- `HashBy[fn]` - used as `Annotated[T, HashBy[fn]]`; hash by `fn(value)`.
|
|
340
|
+
- `NoHash[T]` - exclude a value from hashing (alias for `Annotated[T, HashBy[to_none]]`).
|
|
341
|
+
- `CaptureMe[T]` - capture a global into the call hash on every call.
|
|
342
|
+
- `CaptureMeOnce[T]` - capture a global once per session.
|
|
343
|
+
- `AwaitableValue` - the internal wrapper for async results. You normally never touch it; reach for `set_awaitable` instead of constructing one by hand.
|
|
344
|
+
- `CachedFunction`, `Checkpointer`, `FunctionIdent` - core types.
|
|
345
|
+
- `CheckpointError` - raised by `get` when no valid cache exists.
|
|
346
|
+
- `Storage`, `PickleStorage`, `MemoryStorage` - storage backends.
|
|
347
|
+
- `ObjectHash` - the hashing engine (handles arbitrary objects, NumPy/PyTorch arrays, circular references, and `__objecthash__`).
|
|
348
|
+
|
|
349
|
+
## Pre-configured checkpointers
|
|
350
|
+
|
|
351
|
+
Ready-made presets, importable from `checkpointer`:
|
|
352
|
+
|
|
353
|
+
| Name | Equivalent to |
|
|
354
|
+
| --- | --- |
|
|
355
|
+
| `checkpoint` | `Checkpointer()` - disk-backed defaults. |
|
|
356
|
+
| `capture_checkpoint` | `Checkpointer(capture=True)` - captures all referenced globals. |
|
|
357
|
+
| `memory_checkpoint` | `Checkpointer(storage="memory", verbosity=0)` - in-process, silent. |
|
|
358
|
+
| `tmp_checkpoint` | `Checkpointer(directory="<tmp>/checkpoints")` - stored in the system temp dir. |
|
|
359
|
+
| `static_checkpoint` | `Checkpointer(fn_hash_from=())` - disables code-aware invalidation; the identity hash is fixed until you change `fn_hash_from`. |
|
|
360
|
+
|
|
361
|
+
## Module-level functions
|
|
362
|
+
|
|
363
|
+
- `cleanup_all(invalidated=True, expired=True)` - run `cleanup` on every live `CachedFunction`.
|
|
364
|
+
- `cleanup_memory_storage()` - drop in-memory checkpoints for functions that no longer exist.
|
|
365
|
+
- `get_function_hash(fn)` - compute a function's identity hash without decorating it.
|
|
366
|
+
|
|
367
|
+
## Custom `Storage` interface
|
|
368
|
+
|
|
369
|
+
Subclass `Storage` and implement the methods below. The base class provides `fn_id()`, `fn_dir()`, and expiry helpers.
|
|
370
|
+
|
|
371
|
+
```python
|
|
372
|
+
class Storage:
|
|
373
|
+
checkpointer: Checkpointer
|
|
374
|
+
cached_fn: CachedFunction
|
|
375
|
+
|
|
376
|
+
def store(self, call_hash, data) -> Any: ... # persist & return data
|
|
377
|
+
def exists(self, call_hash) -> bool: ...
|
|
378
|
+
def load(self, call_hash) -> Any: ...
|
|
379
|
+
def delete(self, call_hash) -> None: ...
|
|
380
|
+
def checkpoint_date(self, call_hash) -> datetime: ...
|
|
381
|
+
def cleanup(self, invalidated=True, expired=True) -> None: ...
|
|
382
|
+
def clear(self) -> None: ...
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
The `"pickle"` backend additionally serializes Polars `DataFrame`s as Parquet when Polars is installed.
|