litdata 0.2.66__tar.gz → 0.2.68__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {litdata-0.2.66 → litdata-0.2.68}/CONTRIBUTING.md +1 -1
- {litdata-0.2.66/src/litdata.egg-info → litdata-0.2.68}/PKG-INFO +259 -100
- {litdata-0.2.66 → litdata-0.2.68}/README.md +257 -99
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__about__.py +1 -1
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__init__.py +6 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/constants.py +6 -0
- litdata-0.2.68/src/litdata/debugger.py +397 -0
- litdata-0.2.68/src/litdata/exceptions.py +36 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/data_processor.py +226 -26
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/functions.py +53 -6
- litdata-0.2.68/src/litdata/raw/dataset.py +1712 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/indexer.py +63 -19
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/async_prefetch.py +10 -2
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/cache.py +2 -2
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/client.py +35 -10
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/combined.py +3 -12
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/compression.py +28 -8
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/config.py +9 -17
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/dataloader.py +24 -3
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/dataset.py +108 -9
- litdata-0.2.68/src/litdata/streaming/dataset_update.py +299 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/downloader.py +276 -132
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/item_loader.py +317 -118
- litdata-0.2.68/src/litdata/streaming/posix_fast.py +396 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/reader.py +117 -36
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/resolver.py +16 -2
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/shuffle.py +52 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/writer.py +26 -13
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/format.py +1 -1
- litdata-0.2.68/src/litdata/utilities/keys_index.py +868 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/shuffle.py +114 -0
- {litdata-0.2.66 → litdata-0.2.68/src/litdata.egg-info}/PKG-INFO +259 -100
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/SOURCES.txt +4 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/requires.txt +3 -0
- litdata-0.2.66/src/litdata/debugger.py +0 -205
- litdata-0.2.66/src/litdata/raw/dataset.py +0 -223
- {litdata-0.2.66 → litdata-0.2.68}/LICENSE +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/MANIFEST.in +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/requirements.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/setup.cfg +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/setup.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__main__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/commands.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/cache.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/optimize.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/parser.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/helpers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/imports.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/readers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/utilities.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/types.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/requirements.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/fs_provider.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/parallel.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/sampler.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/serializers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/timing.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/_pytree.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/base.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/breakpoint.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/broadcast.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/dataset_utilities.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/encryption.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/env.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/hf_dataset.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/packing.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/parquet.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/subsample.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/torch_utils.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/train_test_split.py +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/dependency_links.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/entry_points.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/not-zip-safe +0 -0
- {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/top_level.txt +0 -0
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Welcome to the PyTorch Lightning community! We're building the most advanced research platform on the planet to implement the latest, best practices and integrations that the amazing PyTorch team and other research organization rolls out!
|
|
4
4
|
|
|
5
|
-
If you are new to open source, check out [
|
|
5
|
+
If you are new to open source, check out [GitHub's guide to making your first contribution](https://docs.github.com/en/get-started/quickstart/contributing-to-projects).
|
|
6
6
|
|
|
7
7
|
## Main Core Value: One less thing to remember
|
|
8
8
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: litdata
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.68
|
|
4
4
|
Summary: The Deep Learning framework to train, deploy, and ship AI products Lightning fast.
|
|
5
5
|
Home-page: https://github.com/Lightning-AI/litdata
|
|
6
6
|
Download-URL: https://github.com/Lightning-AI/litdata
|
|
@@ -44,6 +44,7 @@ Requires-Dist: pillow; extra == "extras"
|
|
|
44
44
|
Requires-Dist: polars; extra == "extras"
|
|
45
45
|
Requires-Dist: pyarrow<25.0.0; extra == "extras"
|
|
46
46
|
Requires-Dist: tqdm; extra == "extras"
|
|
47
|
+
Requires-Dist: uvloop; sys_platform != "win32" and extra == "extras"
|
|
47
48
|
Requires-Dist: viztracer; extra == "extras"
|
|
48
49
|
Dynamic: author
|
|
49
50
|
Dynamic: author-email
|
|
@@ -71,12 +72,13 @@ Dynamic: summary
|
|
|
71
72
|
|
|
72
73
|
|
|
73
74
|
<pre>
|
|
74
|
-
Transform Optimize
|
|
75
|
+
Transform Optimize / Stream
|
|
75
76
|
|
|
76
|
-
✅ Parallelize data processing ✅ Stream
|
|
77
|
-
✅ Create vector embeddings ✅
|
|
78
|
-
✅ Run distributed inference ✅
|
|
79
|
-
✅ Scrape websites at scale ✅
|
|
77
|
+
✅ Parallelize data processing ✅ Stream raw files with no prep
|
|
78
|
+
✅ Create vector embeddings ✅ Stream large cloud datasets
|
|
79
|
+
✅ Run distributed inference ✅ Accelerate training by 20x
|
|
80
|
+
✅ Scrape websites at scale ✅ Pause and resume data streaming
|
|
81
|
+
✅ Use remote data without local loading
|
|
80
82
|
</pre>
|
|
81
83
|
|
|
82
84
|
---
|
|
@@ -92,6 +94,7 @@ Transform Optimize
|
|
|
92
94
|
<a href="#speed-up-model-training">Optimize data</a> •
|
|
93
95
|
<a href="#transform-datasets">Transform data</a> •
|
|
94
96
|
<a href="#key-features">Features</a> •
|
|
97
|
+
<a href="#stream-raw">Stream raw files</a> •
|
|
95
98
|
<a href="#resolve-paths">Paths & cloud URLs</a> •
|
|
96
99
|
<a href="#benchmarks">Benchmarks</a> •
|
|
97
100
|
<a href="#start-from-a-template">Templates</a> •
|
|
@@ -146,6 +149,8 @@ Install all the extras
|
|
|
146
149
|
pip install 'litdata[extras]'
|
|
147
150
|
```
|
|
148
151
|
|
|
152
|
+
On Linux/macOS, `[extras]` includes optional `uvloop` for a faster asyncio event loop used by `StreamingRawDataset` (stdlib asyncio is the fallback when it is not installed).
|
|
153
|
+
|
|
149
154
|
</details>
|
|
150
155
|
|
|
151
156
|
<details>
|
|
@@ -168,29 +173,38 @@ Source: [`.claude/skills/litdata/`](.claude/skills/litdata/) in this repository
|
|
|
168
173
|
# Speed up model training
|
|
169
174
|
Stream datasets directly from cloud storage without local downloads. Choose the approach that fits your workflow:
|
|
170
175
|
|
|
171
|
-
## Option 1:
|
|
172
|
-
|
|
176
|
+
## Option 1: Stream existing files as-is ⚡⚡ — `StreamingRawDataset`
|
|
177
|
+
|
|
178
|
+
**No optimize step.** Point LitData at a folder of images, audio, text, or any files (local or cloud) and train with a normal PyTorch `DataLoader`. Downloads are **fully asynchronous** and **batched**; cloud clients include **built-in retries**. You receive **raw `bytes`** — decode, parse, or transform however you want.
|
|
179
|
+
|
|
180
|
+
Details → [Stream raw files](#stream-raw).
|
|
173
181
|
|
|
174
182
|
```python
|
|
175
183
|
from litdata import StreamingRawDataset
|
|
176
184
|
from torch.utils.data import DataLoader
|
|
185
|
+
from PIL import Image
|
|
186
|
+
import io
|
|
177
187
|
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
188
|
+
dataset = StreamingRawDataset(
|
|
189
|
+
"s3://my-bucket/raw-images/", # or gs://, azure://, /teamspace/s3_connections/..., local path
|
|
190
|
+
transform=lambda b: Image.open(io.BytesIO(b)).convert("RGB"), # optional — default is raw bytes
|
|
191
|
+
)
|
|
192
|
+
loader = DataLoader(dataset, batch_size=32, num_workers=8)
|
|
181
193
|
|
|
182
|
-
for batch in
|
|
183
|
-
|
|
184
|
-
pass
|
|
194
|
+
for batch in loader:
|
|
195
|
+
train_step(batch)
|
|
185
196
|
```
|
|
186
197
|
|
|
187
198
|
**Key benefits:**
|
|
188
199
|
|
|
189
|
-
✅ **
|
|
190
|
-
✅ **
|
|
191
|
-
✅ **
|
|
192
|
-
✅ **
|
|
193
|
-
✅ **Cloud-native:**
|
|
200
|
+
✅ **Zero preprocess:** No chunking job — use the files you already have.
|
|
201
|
+
✅ **Raw bytes, your rules:** Each sample is file `bytes`; decode with PIL, torchaudio, json, or any custom logic (`transform=` optional).
|
|
202
|
+
✅ **Fully async + batched:** Concurrent downloads via `asyncio` / `__getitems__` (not one-file-at-a-time).
|
|
203
|
+
✅ **Built-in retries:** Cloud downloads retry transient failures (adaptive client retries).
|
|
204
|
+
✅ **Cloud-native:** S3 / GCS / Azure / Studio connections; same path resolver as optimized streaming.
|
|
205
|
+
✅ **Grouped samples:** Override `setup()` to yield image+mask, audio+transcript, etc.
|
|
206
|
+
✅ **Indexed once:** `index.json.zstd` cached locally and on the bucket for fast restarts.
|
|
207
|
+
✅ **Upgrade path:** When I/O becomes the bottleneck, `optimize` → `StreamingDataset` for max throughput.
|
|
194
208
|
|
|
195
209
|
## Option 2: Optimize for maximum performance ⚡⚡⚡
|
|
196
210
|
Accelerate model training (20x faster) by optimizing datasets for streaming directly from cloud storage. Work with remote data without local downloads with features like loading data subsets, accessing individual samples, and resumable streaming.
|
|
@@ -227,7 +241,7 @@ if __name__ == "__main__":
|
|
|
227
241
|
inputs=list(range(1000)), # the inputs to the function (here it's a list of numbers)
|
|
228
242
|
output_dir="fast_data", # optimized data is stored here
|
|
229
243
|
num_workers=4, # the number of workers on the same machine
|
|
230
|
-
chunk_bytes="64MB" #
|
|
244
|
+
chunk_bytes="64MB" # default; see FAQ for larger samples
|
|
231
245
|
)
|
|
232
246
|
```
|
|
233
247
|
|
|
@@ -265,6 +279,22 @@ for sample in dataloader:
|
|
|
265
279
|
img, cls = sample["image"], sample["class"]
|
|
266
280
|
```
|
|
267
281
|
|
|
282
|
+
**Keyed lookup and in-place patches** (needs `polars` and `optimize(..., key_fn=...)` or `build_keys_index`):
|
|
283
|
+
|
|
284
|
+
```python
|
|
285
|
+
ld.optimize(fn=fn, inputs=inputs, output_dir="fast_data", chunk_bytes="64MB", key_fn=lambda s: s["id"])
|
|
286
|
+
|
|
287
|
+
ds = ld.StreamingDataset("fast_data")
|
|
288
|
+
sample = ds["entity-id"] # str keys
|
|
289
|
+
sample = ds.get_by_key(42) # int entity keys; ds[42] is still positional
|
|
290
|
+
|
|
291
|
+
with ld.dataset_update("fast_data") as update: # local directory only
|
|
292
|
+
update["entity-id"] = {"id": "entity-id", "x": 1}
|
|
293
|
+
update.commit()
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
`mode="append"` continues chunk numbering. `use_checkpoint=True` tries to resume an interrupted optimize. They are not the same.
|
|
297
|
+
|
|
268
298
|
**Key benefits:**
|
|
269
299
|
|
|
270
300
|
✅ **Accelerate training:** Optimized datasets load 20x faster.
|
|
@@ -325,79 +355,158 @@ ld.map(
|
|
|
325
355
|
## Features for optimizing and streaming datasets for model training
|
|
326
356
|
|
|
327
357
|
<details>
|
|
328
|
-
<summary> ✅ Stream raw
|
|
358
|
+
<summary> ✅ Stream raw files as-is (no optimize) — StreamingRawDataset <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
|
|
329
359
|
|
|
330
360
|
|
|
331
|
-
|
|
361
|
+
`StreamingRawDataset` streams **your existing files** from local disk or cloud storage with **no conversion step**. It is a map-style `torch.utils.data.Dataset`: use a standard PyTorch `DataLoader` (not `StreamingDataLoader`).
|
|
332
362
|
|
|
333
|
-
**
|
|
363
|
+
**You get raw `bytes`.** LitData does not impose a sample schema — open images with PIL, parse JSONL, decode audio, run your own tokenizer, or pass a `transform=` if you prefer. Grouped items yield `list[bytes]` (e.g. image + mask).
|
|
334
364
|
|
|
335
|
-
|
|
365
|
+
Downloads are **fully asynchronous** and **batched**: when the DataLoader requests a batch, `__getitems__` fetches those files concurrently with `asyncio.gather`. Cloud clients include **built-in retries** for transient network errors.
|
|
336
366
|
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
367
|
+
Use it when you want to train or prototype on JPEGs, masks, audio, JSONL, etc. **immediately**. Switch to [`optimize` → `StreamingDataset`](#speed-up-model-training) later if you need maximum cloud training throughput.
|
|
368
|
+
|
|
369
|
+
| | `StreamingRawDataset` | `StreamingDataset` (optimized) |
|
|
370
|
+
|--|----------------------|--------------------------------|
|
|
371
|
+
| Prep | None — point at a folder | One-time `optimize` → `chunk-*.bin` + `index.json` |
|
|
372
|
+
| Item | **Raw file `bytes`** (you decide how to decode) | Deserialized samples (dict/tensor/…) |
|
|
373
|
+
| I/O | Fully async, batched downloads + retries | Chunk prefetch / cache pipeline |
|
|
374
|
+
| Loader | `torch.utils.data.DataLoader` | Prefer `StreamingDataLoader` (shuffle, resume) |
|
|
375
|
+
| Best for | Instant start, full control over bytes | Highest sustained training I/O |
|
|
340
376
|
|
|
341
|
-
|
|
342
|
-
|
|
377
|
+
### Install (cloud)
|
|
378
|
+
|
|
379
|
+
```bash
|
|
380
|
+
pip install "litdata[extra]" s3fs # Amazon S3
|
|
381
|
+
pip install "litdata[extra]" gcsfs # Google Cloud Storage
|
|
382
|
+
# Azure / Studio connections: see Paths & cloud URLs
|
|
343
383
|
```
|
|
344
384
|
|
|
345
|
-
|
|
385
|
+
### Quick start
|
|
386
|
+
|
|
346
387
|
```python
|
|
347
388
|
from torch.utils.data import DataLoader
|
|
348
389
|
from litdata import StreamingRawDataset
|
|
390
|
+
from PIL import Image
|
|
391
|
+
import io
|
|
349
392
|
|
|
350
|
-
|
|
393
|
+
def to_image(data: bytes):
|
|
394
|
+
return Image.open(io.BytesIO(data)).convert("RGB")
|
|
395
|
+
|
|
396
|
+
dataset = StreamingRawDataset(
|
|
397
|
+
"s3://my-bucket/images/", # also: gs://, azure://, /teamspace/s3_connections/..., local path
|
|
398
|
+
transform=to_image, # optional; default yields raw bytes
|
|
399
|
+
storage_options={}, # optional cloud credentials / endpoint
|
|
400
|
+
)
|
|
401
|
+
loader = DataLoader(dataset, batch_size=32, num_workers=8)
|
|
351
402
|
|
|
352
|
-
# Use with PyTorch DataLoader
|
|
353
|
-
loader = DataLoader(dataset, batch_size=32)
|
|
354
403
|
for batch in loader:
|
|
355
|
-
|
|
356
|
-
pass
|
|
404
|
+
train_step(batch)
|
|
357
405
|
```
|
|
358
406
|
|
|
359
|
-
|
|
407
|
+
### Constructor knobs
|
|
408
|
+
|
|
409
|
+
| Arg | Default | Purpose |
|
|
410
|
+
|-----|---------|---------|
|
|
411
|
+
| `input_dir` | required | Folder URL/path (same [resolver](#resolve-paths) as optimized streaming) |
|
|
412
|
+
| `cache_dir` | LitData default cache | Where the file index (and optional file cache) live |
|
|
413
|
+
| `cache_files` | `False` | If `True`, keep downloaded files on disk under `cache_dir` (mirror remote layout) |
|
|
414
|
+
| `recompute_index` | `False` | Force re-scan when remote files changed |
|
|
415
|
+
| `transform` | `None` | `fn(bytes) -> Any` or `fn(list[bytes]) -> Any` for grouped items |
|
|
416
|
+
| `storage_options` | `{}` | Cloud client options |
|
|
417
|
+
| `indexer` | `FileIndexer()` | Custom discovery (subclass `BaseIndexer`) |
|
|
418
|
+
| `max_concurrent_downloads` | `None` (adaptive) | Per-worker in-flight downloads. `None` = size-aware budget (bandwidth; Little’s-law only for medians <~8 MiB) split across workers; single-process capped at 128. An explicit `int` is used exactly (no silent clamp) |
|
|
419
|
+
| `max_prefetch` | `16` | Per-worker sequential look-ahead after each batch (default on). When `num_workers > 1`, effective look-ahead is `min(max_prefetch, 64 // num_workers)` so aggregate stays ~64 items. Pass `0` to disable |
|
|
420
|
+
| `prefetch_cache_size` | auto | LRU cap for prefetched items (defaults from `max_prefetch`) |
|
|
421
|
+
| `hedge_delay` | `0` | Seconds before a hedged duplicate GET for a slow download (`0` = off, default; opt-in) |
|
|
422
|
+
| `range_parallel_threshold` | `0` | Objects ≥ this many bytes use parallel ranged GETs (`0` = whole-object only; opt-in) |
|
|
423
|
+
| `item_type` | `"bytes"` | `"bytes"` buffers in RAM; `"path"` returns local cache paths (`cache_files=True` required) |
|
|
360
424
|
|
|
425
|
+
### Group related files (`setup`)
|
|
361
426
|
|
|
362
|
-
|
|
427
|
+
Default: **one file = one sample**. Override `setup` to filter or group (image + mask, audio + transcript, …). Return either a list of `FileMetadata` or a list of groups (`list[list[FileMetadata]]`).
|
|
363
428
|
|
|
364
429
|
```python
|
|
365
|
-
from
|
|
430
|
+
from collections import defaultdict
|
|
366
431
|
from torch.utils.data import DataLoader
|
|
367
432
|
from litdata import StreamingRawDataset
|
|
368
433
|
from litdata.raw.indexer import FileMetadata
|
|
369
434
|
|
|
370
435
|
class SegmentationRawDataset(StreamingRawDataset):
|
|
371
|
-
def setup(self, files: list[FileMetadata]) ->
|
|
372
|
-
#
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
436
|
+
def setup(self, files: list[FileMetadata]) -> list[list[FileMetadata]]:
|
|
437
|
+
# Pair img_001.jpg with img_001.png (mask) by stem
|
|
438
|
+
by_stem: dict[str, dict[str, FileMetadata]] = defaultdict(dict)
|
|
439
|
+
for f in files:
|
|
440
|
+
name = f.path.rsplit("/", 1)[-1]
|
|
441
|
+
stem, _, ext = name.rpartition(".")
|
|
442
|
+
by_stem[stem][ext.lower()] = f
|
|
443
|
+
items = []
|
|
444
|
+
for stem, parts in sorted(by_stem.items()):
|
|
445
|
+
if "jpg" in parts and "png" in parts:
|
|
446
|
+
items.append([parts["jpg"], parts["png"]])
|
|
447
|
+
return items
|
|
448
|
+
|
|
449
|
+
dataset = SegmentationRawDataset(
|
|
450
|
+
"s3://bucket/seg/",
|
|
451
|
+
transform=lambda pair: (pair[0], pair[1]), # list[bytes]: [image, mask]
|
|
452
|
+
)
|
|
453
|
+
loader = DataLoader(dataset, batch_size=16, num_workers=4)
|
|
454
|
+
for images, masks in loader:
|
|
455
|
+
...
|
|
385
456
|
```
|
|
386
457
|
|
|
387
|
-
|
|
458
|
+
### Index caching (`index.json.zstd`)
|
|
388
459
|
|
|
389
|
-
|
|
460
|
+
First open scans the tree and writes a compressed file list:
|
|
390
461
|
|
|
391
|
-
|
|
392
|
-
- **
|
|
393
|
-
- **Remote:** Automatically saved to cloud storage (e.g., `s3://bucket/files/index.json.zstd`) for reuse
|
|
462
|
+
- **Local cache** under your LitData cache dir (fast restart on the same machine)
|
|
463
|
+
- **Remote copy** next to the data when possible (e.g. `s3://bucket/files/index.json.zstd`) so every machine skips the scan
|
|
394
464
|
|
|
395
|
-
**Force Rebuild:**
|
|
396
465
|
```python
|
|
397
|
-
#
|
|
466
|
+
# After adding/removing files on the bucket:
|
|
398
467
|
dataset = StreamingRawDataset("s3://bucket/files/", recompute_index=True)
|
|
399
468
|
```
|
|
400
469
|
|
|
470
|
+
Do **not** confuse this with optimized LitData’s `index.json` (chunk metadata). Raw indexing only lists files.
|
|
471
|
+
|
|
472
|
+
### How downloads work
|
|
473
|
+
|
|
474
|
+
1. DataLoader asks for a batch of indices → `__getitems__`.
|
|
475
|
+
2. LitData **asynchronously** downloads those files **in parallel** (`asyncio.gather` + `adownload_fileobj`).
|
|
476
|
+
3. Cloud SDKs apply **retries** on transient failures (e.g. S3 adaptive retries).
|
|
477
|
+
4. Each item is returned as **`bytes`** (or `list[bytes]` if `setup` grouped files), then optional `transform`.
|
|
478
|
+
|
|
479
|
+
Your training loop stays normal PyTorch — no async/`await` in user code.
|
|
480
|
+
|
|
481
|
+
```python
|
|
482
|
+
# Default: you own the bytes
|
|
483
|
+
dataset = StreamingRawDataset("s3://bucket/files/")
|
|
484
|
+
raw: bytes = dataset[0]
|
|
485
|
+
# e.g. Image.open(io.BytesIO(raw)), json.loads(raw), np.frombuffer(raw), ...
|
|
486
|
+
```
|
|
487
|
+
|
|
488
|
+
### Tips
|
|
489
|
+
|
|
490
|
+
- Prefer `num_workers > 0` so worker processes overlap async batch downloads with training. Scale workers toward host vCPUs for network-bound JPEG-sized objects — avoid saturating every vCPU.
|
|
491
|
+
- On Linux, after any parent-process dataset I/O, use `DataLoader(..., multiprocessing_context="spawn", persistent_workers=True)` — default `fork` can hang S3 clients in workers.
|
|
492
|
+
- Default `max_prefetch=16` enables sequential look-ahead **per DataLoader worker**; shuffled access disables it. Pass `0` to turn off. When `num_workers > 1`, look-ahead and download concurrency both scale down with worker count so aggregate in-flight work stays bounded.
|
|
493
|
+
- Prefer an `s3://` / `gs://` URL or `/teamspace/s3_connections/...` so LitData hits the bucket directly ([resolver](#resolve-paths)) — avoid reading through FUSE.
|
|
494
|
+
- Leave `range_parallel_threshold=0` (default) for typical JPEGs; raise it only for large objects where parallel ranged GETs help.
|
|
495
|
+
- Best for medium/large files. Tiny objects (≲100 KB) are request-overhead bound — pack with [`optimize`](#speed-up-model-training) → `StreamingDataset` when I/O plateaus.
|
|
496
|
+
|
|
497
|
+
### Throughput
|
|
498
|
+
|
|
499
|
+
On ImageNet val raw over S3 (50 k JPEGs, batch size 64, spawn workers), throughput gains are clearest at **low worker counts / notebooks** (**+20–80%** at ≤8 workers). At **high workers** (≥16), results are roughly **parity within run-to-run noise**.
|
|
500
|
+
|
|
501
|
+
| workers | before | after | Δ |
|
|
502
|
+
|--------:|-------:|------:|--:|
|
|
503
|
+
| 0 | 543 | 735 | **+35%** |
|
|
504
|
+
| 2 | 816 | 1475 | **+81%** |
|
|
505
|
+
| 8 | 4841 | 5718 | **+18%** |
|
|
506
|
+
| 16+ | ~6k | ~6k | ~parity |
|
|
507
|
+
|
|
508
|
+
Useful knobs: `num_workers`, `max_prefetch` (default 16; worker-aware), `download_timeout` (batch-level hang protection). Ranged parallel downloads stay opt-in (`range_parallel_threshold=0`).
|
|
509
|
+
|
|
401
510
|
</details>
|
|
402
511
|
|
|
403
512
|
<details>
|
|
@@ -692,6 +801,10 @@ Shuffling is **deterministic** and designed for distributed training:
|
|
|
692
801
|
|
|
693
802
|
The permutation depends on `seed`, the epoch, and chunk metadata — the same settings always yield the same order (required for resumable `state_dict`).
|
|
694
803
|
|
|
804
|
+
**Object storage (`s3://`, `gs://`, …)** globally permutes chunks (`FullShuffle`). Random chunk order is cheap once files are already copied into the local cache.
|
|
805
|
+
|
|
806
|
+
**POSIX-fast** (automatic for any local path) mmaps chunks in place. **Vast / NFS / Lustre / GPFS** (and `LITDATA_POSIX_FAST=1`) use `WindowShuffle`: each worker gets **whole chunks** in a sequential stripe, then shuffles only inside a sliding window (default **16**, `LITDATA_POSIX_SHUFFLE_WINDOW`) for both chunk order and in-chunk items. Local disks (ext4/xfs) keep global `FullShuffle`. Object URLs stay on `FullShuffle`. `LITDATA_POSIX_FAST=0` disables in-place mmap.
|
|
807
|
+
|
|
695
808
|
```python
|
|
696
809
|
from litdata import StreamingDataset, StreamingDataLoader
|
|
697
810
|
|
|
@@ -715,6 +828,35 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
|
|
|
715
828
|
|
|
716
829
|
</details>
|
|
717
830
|
|
|
831
|
+
<details>
|
|
832
|
+
<summary> ✅ FAQ: chunk size & shuffle before optimize <a id="faq-chunk-shuffle" href="#faq-chunk-shuffle">🔗</a> </summary>
|
|
833
|
+
|
|
834
|
+
|
|
835
|
+
### What `chunk_bytes` should I use?
|
|
836
|
+
|
|
837
|
+
Default is **64MB** — a good starting point for typical small/medium samples.
|
|
838
|
+
|
|
839
|
+
When each datapoint is large (e.g. a few MB), prefer a **larger chunk** (practical range often **256–512MB**) so each chunk holds more samples and **intra-chunk batch randomization** has a bigger pool. Tradeoff: larger chunks take **longer to download** before they can be used.
|
|
840
|
+
|
|
841
|
+
This is expert guidance (recommended-range mindset), not a published chunk-size sweep.
|
|
842
|
+
|
|
843
|
+
### Is StreamingDataset shuffle enough if my source data is ordered?
|
|
844
|
+
|
|
845
|
+
**Not always.** LitData handles **distributed sampling** and **bucket sampling within chunks** automatically (`shuffle=True` randomizes chunk order and item order inside each chunk). That is **not** a substitute for a fully shuffled file-level DataLoader when the source has strong structure (same subject/set contiguous, class blocks, etc.).
|
|
846
|
+
|
|
847
|
+
If ordered data would make chunked sampling problematic and you cannot embed the grouping as the sample unit:
|
|
848
|
+
|
|
849
|
+
- Shuffle the list of samples **before** `optimize` so chunks mix well, **or**
|
|
850
|
+
- Use [`StreamingRawDataset`](#stream-raw) (per-file random access via a standard PyTorch `DataLoader` with `shuffle=True`) instead of optimize → `StreamingDataset`.
|
|
851
|
+
|
|
852
|
+
### FUSE vs LitData (Lightning Studios)
|
|
853
|
+
|
|
854
|
+
`/teamspace/s3_connections` (and related mounts) are **FUSE** — fine for browsing, not for training I/O. Under load they are very slow and can crash. Pass the same path into LitData (`StreamingRawDataset` / `StreamingDataset` / `optimize`): LitData resolves it and talks **directly** to the bucket ([Resolve any path](#resolve-paths)).
|
|
855
|
+
|
|
856
|
+
Rough ImageNet order-of-magnitude on a Studio (not hard guarantees; right tuning for raw): FUSE hand-read ~**600** images/s · [`StreamingRawDataset`](#stream-raw) ~**6–7k** · optimized [`StreamingDataset`](#speed-up-model-training) (64MB chunks) ~**11k**.
|
|
857
|
+
|
|
858
|
+
</details>
|
|
859
|
+
|
|
718
860
|
<details>
|
|
719
861
|
<summary> ✅ StreamingDataset & StreamingDataLoader knobs <a id="streaming-kwargs" href="#streaming-kwargs">🔗</a> </summary>
|
|
720
862
|
|
|
@@ -731,7 +873,7 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
|
|
|
731
873
|
| `seed` | `42` | Shuffle / subsample RNG |
|
|
732
874
|
| `serializers` | built-ins | Custom serialize/deserialize map |
|
|
733
875
|
| `max_cache_size` | `"100GB"` | Evict consumed chunks beyond this size |
|
|
734
|
-
| `max_pre_download` | `2` | Chunks each worker may prefetch (raise for throughput; watch disk) |
|
|
876
|
+
| `max_pre_download` | `2` | Chunks each worker may prefetch (raise for throughput; watch disk / RAM) |
|
|
735
877
|
| `subsample` | `1.0` | Fraction of data (`0.01`) or upsample (`2.5`) |
|
|
736
878
|
| `encryption` | `None` | `FernetEncryption` / `RSAEncryption` / custom |
|
|
737
879
|
| `storage_options` | `{}` | Cloud client options |
|
|
@@ -742,6 +884,8 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
|
|
|
742
884
|
|
|
743
885
|
Peak disk ≈ `num_workers × max_pre_download × mean_chunk_size`.
|
|
744
886
|
|
|
887
|
+
On **Vast / NFS / local disk**, POSIX-fast is on by default (`LITDATA_POSIX_FAST=0` to disable). `WILLNEED` prefetch and `num_workers` are capped when they would exceed about half of `MemAvailable`. Idle **hugepages** (common on GPU nodes) do not count as available RAM — drop unused `nr_hugepages` if `MemAvailable` looks tiny next to `MemTotal`.
|
|
888
|
+
|
|
745
889
|
**`StreamingDataLoader`**
|
|
746
890
|
|
|
747
891
|
| Argument | Description |
|
|
@@ -1093,7 +1237,7 @@ if __name__ == "__main__":
|
|
|
1093
1237
|
|
|
1094
1238
|
Mix and match different sets of data to experiment and create better models.
|
|
1095
1239
|
|
|
1096
|
-
Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://
|
|
1240
|
+
Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
|
|
1097
1241
|
|
|
1098
1242
|
```python
|
|
1099
1243
|
from litdata import StreamingDataset, CombinedStreamingDataset, StreamingDataLoader, TokensLoader
|
|
@@ -1684,7 +1828,7 @@ Only **worker 0** is instrumented. When an `int` is used, the tracer wraps `fetc
|
|
|
1684
1828
|
|
|
1685
1829
|
- Delete or change `profile_dir` between runs — LitData removes an existing `result.json` before starting.
|
|
1686
1830
|
- Pair with a wiped chunk cache if you care about **cold** epoch behavior (`litdata cache clear`).
|
|
1687
|
-
- For deeper LitData internals (download /
|
|
1831
|
+
- For deeper LitData internals (download / read / delete timeline), use `enable_tracer()` + [Litracer](https://github.com/Lightning-AI/litracer) instead — see [Debug & Profile LitData](#debug-profile). That path is complementary: viztracer = DataLoader worker CPU timeline; Litracer = LitData pipeline events.
|
|
1688
1832
|
|
|
1689
1833
|
</details>
|
|
1690
1834
|
|
|
@@ -1797,6 +1941,9 @@ export LITDATA_ASYNC_MIN_PRE_DOWNLOAD=0
|
|
|
1797
1941
|
| `LITDATA_DISABLE_VERSION_CHECK` | `0` | `1` skips the upgrade tip |
|
|
1798
1942
|
| `HF_TOKEN` | — | Gated Hugging Face datasets |
|
|
1799
1943
|
| `DEBUG_LITDATA` / `PRINT_DEBUG_LOGS` | `0` | Internal debug / stdout logs |
|
|
1944
|
+
| `LITDATA_LOG_FILE` | `litdata_debug.log` | `enable_tracer()` output path |
|
|
1945
|
+
| `LITDATA_TRACE_LEVEL` | unset | `batch` / `chunk` / `sample` / `debug` / `off` (see [Debug & Profile](#debug-profile)) |
|
|
1946
|
+
| `LITDATA_TRACE_CATEGORIES` | from level | Comma-separated cats, e.g. `download,read,delete` |
|
|
1800
1947
|
|
|
1801
1948
|
Multi-node `optimize`/`map` on Studios also uses `DATA_OPTIMIZER_*` (set by the platform). Full catalog (debug logs, Studio injects, torchrun): see the LitData skill `reference/env-vars.md` when using agent skills, or the source modules `constants.py` / `async_prefetch.py`.
|
|
1802
1949
|
|
|
@@ -1962,66 +2109,68 @@ ds = StreamingDataset(input_dir=data_dir, encryption=rsa)
|
|
|
1962
2109
|
|
|
1963
2110
|
|
|
1964
2111
|
|
|
1965
|
-
|
|
2112
|
+
`enable_tracer()` records the streaming pipeline (download vs read vs delete vs batch) as one-line events. [Litracer](https://github.com/Lightning-AI/litracer) converts that log into a Chrome / [Perfetto](https://ui.perfetto.dev) trace.
|
|
1966
2113
|
|
|
1967
|
-
|
|
2114
|
+
This is complementary to [`profile_batches`](#profile-loading) (viztracer = DataLoader worker **CPU**; Litracer = LitData **pipeline** events).
|
|
1968
2115
|
|
|
1969
|
-
-
|
|
2116
|
+
<img width="1439" alt="431247797-0e955e71-2f9a-4aad-b7c1-a8218fed2e2e" src="https://github.com/user-attachments/assets/4e40676c-ba0b-49af-acac-975977173669" />
|
|
1970
2117
|
|
|
1971
2118
|
```python
|
|
1972
2119
|
import litdata as ld
|
|
1973
2120
|
from litdata.debugger import enable_tracer
|
|
1974
2121
|
|
|
1975
|
-
#
|
|
1976
|
-
enable_tracer()
|
|
2122
|
+
# Call once per process, before the DataLoader. Delete an existing log before re-tracing (append).
|
|
2123
|
+
enable_tracer(level="chunk", log_file="litdata_debug.log")
|
|
2124
|
+
# level="batch" | "chunk" (default) | "sample" | "debug" | "off"
|
|
2125
|
+
# enable_tracer(categories=["download", "read", "delete"])
|
|
1977
2126
|
|
|
1978
2127
|
if __name__ == "__main__":
|
|
1979
2128
|
dataset = ld.StreamingDataset("s3://my-bucket/my-data", shuffle=True)
|
|
1980
|
-
|
|
1981
|
-
|
|
1982
|
-
for batch in dataloader:
|
|
1983
|
-
print(batch) # Replace with your data processing logic
|
|
2129
|
+
for batch in ld.StreamingDataLoader(dataset, batch_size=64, num_workers=8):
|
|
2130
|
+
...
|
|
1984
2131
|
```
|
|
1985
2132
|
|
|
1986
|
-
|
|
2133
|
+
| Level | Events |
|
|
2134
|
+
| ----- | ------ |
|
|
2135
|
+
| `batch` | Epoch + per-batch spans, plus crashes |
|
|
2136
|
+
| `chunk` (default) | + `download`, `read`, `delete`, `decompress`, `prefetch` |
|
|
2137
|
+
| `sample` | + per-item `__getitem__` (high volume) |
|
|
2138
|
+
| `debug` | + `.cnt` lock refcount spans |
|
|
2139
|
+
| `off` | Disable |
|
|
1987
2140
|
|
|
1988
|
-
|
|
2141
|
+
Event **names** are stable (`download`, `read`, `delete`, `batch`, `sample`, `crash`). Chunk / sample indexes live in args so Perfetto groups all downloads together. Each line is `key: value;` pairs with Chrome **microsecond** timestamps. Crashes are a one-line instant (`ph: I`, `name: crash`); the Python traceback is printed to **stderr**, not the log file (a multi-line `logger.exception` would break Litracer).
|
|
1989
2142
|
|
|
1990
|
-
|
|
1991
|
-
python main.py
|
|
1992
|
-
```
|
|
2143
|
+
Env overrides: `LITDATA_LOG_FILE`, `LITDATA_TRACE_LEVEL`, `LITDATA_TRACE_CATEGORIES` (comma-separated). Tracer calls are no-ops when tracing is off.
|
|
1993
2144
|
|
|
1994
|
-
|
|
2145
|
+
1. Generate the log:
|
|
1995
2146
|
|
|
1996
|
-
|
|
1997
|
-
|
|
1998
|
-
|
|
2147
|
+
```bash
|
|
2148
|
+
python train.py # writes litdata_debug.log
|
|
2149
|
+
```
|
|
1999
2150
|
|
|
2000
|
-
|
|
2001
|
-
go install github.com/deependujha/litracer@latest
|
|
2002
|
-
```
|
|
2151
|
+
2. Install [Litracer](https://github.com/Lightning-AI/litracer) (Go 1.23+):
|
|
2003
2152
|
|
|
2004
|
-
|
|
2005
|
-
|
|
2006
|
-
|
|
2153
|
+
```bash
|
|
2154
|
+
git clone https://github.com/Lightning-AI/litracer.git
|
|
2155
|
+
cd litracer && go build -o litracer .
|
|
2156
|
+
```
|
|
2007
2157
|
|
|
2008
|
-
|
|
2158
|
+
Or `go install github.com/deependujha/litracer@latest` (published Go module path). Until `go.mod` is renamed, `go install github.com/Lightning-AI/litracer@latest` does not work. Release binaries: [GitHub Releases](https://github.com/Lightning-AI/litracer/releases).
|
|
2009
2159
|
|
|
2010
|
-
|
|
2160
|
+
3. Convert and open in Perfetto:
|
|
2011
2161
|
|
|
2012
2162
|
```bash
|
|
2013
|
-
|
|
2163
|
+
litracer --quiet --validate -o litdata_trace.json.gz litdata_debug.log
|
|
2164
|
+
litracer --quiet --cat download,read,delete -o io.json.gz litdata_debug.log
|
|
2165
|
+
# open the .json.gz at https://ui.perfetto.dev (preferred) or chrome://tracing
|
|
2014
2166
|
```
|
|
2015
2167
|
|
|
2016
|
-
|
|
2017
|
-
|
|
2018
|
-
- Use either `chrome://tracing` in the Chrome browser or `ui.perfetto.dev` to view the `litdata_trace.json` file for in-depth performance insights. You can also use `SQL queries` to analyze the logs.
|
|
2019
|
-
- `Perfetto` is recommended over `chrome://tracing` for visualization & analyzing.
|
|
2168
|
+
`--quiet` prints a one-line summary (per-category durations, unmatched B/E, crashes). `--cat` keeps only those categories. Matched B/E pairs become complete (`ph: X`) spans unless `--no-complete`. Default output is gzip Chrome JSON (`.json.gz`) — both Perfetto and `chrome://tracing` open it; pass `-o file.json` for uncompressed.
|
|
2020
2169
|
|
|
2021
|
-
-
|
|
2170
|
+
- For trace files `> 2GB`, see [Perfetto large traces](https://perfetto.dev/docs/visualization/large-traces).
|
|
2171
|
+
- If you connect Perfetto to the RPC server, prefer Chrome over Brave (Brave often does not autodetect the RPC server).
|
|
2022
2172
|
|
|
2023
|
-
|
|
2024
|
-
- If you are trying to connect Perfetto to the RPC server, it is recommended to use Chrome over Brave, as it has been observed that Perfetto in Brave does not autodetect the RPC server.
|
|
2173
|
+
**Multi-worker `s3://` `FileNotFoundError` after ~120s:** `num_workers=0` working while `num_workers>0` fails usually means the DataLoader parent started obstore (tokio) before fork and worker GETs hung. Current LitData fetches `index.json` with boto3 so workers can lazy-init obstore; they fall back to boto3 if the parent already started the runtime. On Studio R2 / `lightning_storage`, the same symptom can be a prefetch-thread crash (`data_connection_id` / `endpoint_url` into `boto3.Session`) — look for `[litdata] PrepareChunksThread CRASHED` on stderr and a `crash` instant in the trace.
|
|
2025
2174
|
|
|
2026
2175
|
</details>
|
|
2027
2176
|
|
|
@@ -2213,7 +2362,7 @@ Full knob list for `litdata.optimize` (see Quick start for the minimal recipe).
|
|
|
2213
2362
|
| `output_dir` | `"optimized_data"` | Local or cloud ([resolver](#resolve-paths)); version remote prefixes |
|
|
2214
2363
|
| `input_dir` | `None` | Remote input root for background download |
|
|
2215
2364
|
| `weights` | `None` | Per-input weights to balance workers |
|
|
2216
|
-
| `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"
|
|
2365
|
+
| `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`; see [FAQ](#faq-chunk-shuffle) for larger samples) |
|
|
2217
2366
|
| `chunk_size` | `None` | Max items (or tokens with `TokensLoader`) per chunk |
|
|
2218
2367
|
| `align_chunking` | `False` | Match single-worker chunk boundaries (needs `chunk_size`; uneven load) |
|
|
2219
2368
|
| `compression` | `None` | `"zstd"` today |
|
|
@@ -2306,6 +2455,17 @@ Speed to stream Imagenet 1.2M from local disk with ffcv vs LitData:
|
|
|
2306
2455
|
| ffcv(os_cache=True) | JPEG 90% | 20 GB | 7653 | 8051 |
|
|
2307
2456
|
| ffcv(os_cache=False) | JPEG 90% | 20 GB | 8149 | 8607 |
|
|
2308
2457
|
|
|
2458
|
+
Speed to stream a **synthetic ImageNet-scale set from Vast NFS** (NFSv3 `nconnect=32`, 208-CPU host, ~1 TiB RAM). Dataset: **1.08M** JPEG q95 256×256 (~160 GiB, 64 MiB chunks). `StreamingDataLoader`, batch **256**, `shuffle=True`, `drop_last=True`, decode only unless noted. POSIX-fast mmaps chunks **in place** (no copy into `~/.lightning/chunks`).
|
|
2459
|
+
|
|
2460
|
+
| Setup | Workers | Images / sec |
|
|
2461
|
+
|---|---|---|
|
|
2462
|
+
| Copy into local cache (`LITDATA_POSIX_FAST=0`) | 48 | **16.7k** (2-epoch avg) |
|
|
2463
|
+
| POSIX-fast (this default on local/Vast paths) | 48 | **18.2k** |
|
|
2464
|
+
| POSIX-fast + README ImageNet augs (crop 224, flip, float32) | 48 | **12.9k** |
|
|
2465
|
+
| POSIX-fast, all CPU cores | **208** | **35.8k** |
|
|
2466
|
+
|
|
2467
|
+
Notes: 208 workers need enough **MemAvailable**. This host had **928×1 GiB hugepages** reserved and idle (~900 GiB locked); after `nr_hugepages=0`, 208 workers stayed healthy. If `num_workers=os.cpu_count()` would crowd RAM, LitData **clamps** workers (`LITDATA_POSIX_MAX_WORKERS=0` disables) and skips `WILLNEED` prefetch. Real ImageNet JPEG 90% is much smaller (~12 GiB) and usually decodes faster than this q95 noise set.
|
|
2468
|
+
|
|
2309
2469
|
### Raw Dataset
|
|
2310
2470
|
|
|
2311
2471
|
Speed to stream raw Imagenet 1.2M from different cloud storage providers:
|
|
@@ -2316,8 +2476,7 @@ Speed to stream raw Imagenet 1.2M from different cloud storage providers:
|
|
|
2316
2476
|
| AWS S3 | ~6400 +/- 100 | ~3200 +/- 100 |
|
|
2317
2477
|
| Google Cloud Storage | ~5650 +/- 100 | ~3100 +/- 100 |
|
|
2318
2478
|
|
|
2319
|
-
> **
|
|
2320
|
-
> Use `StreamingRawDataset` if you want to stream your data as-is. Use `StreamingDataset` if you want the fastest streaming and are okay with optimizing your data first.
|
|
2479
|
+
> **Also see:** [`StreamingRawDataset`](#stream-raw) streams existing files with **no optimize step** (great default to start). Use `StreamingDataset` after `optimize` when you need the highest sustained training throughput.
|
|
2321
2480
|
|
|
2322
2481
|
|
|
2323
2482
|
|
|
@@ -2406,7 +2565,7 @@ Below are templates for real-world applications of LitData at scale.
|
|
|
2406
2565
|
| -------------------------------- | ----------------- | ----------------- | -------------- | -------------- |
|
|
2407
2566
|
| [Benchmark cloud data-loading libraries](https://lightning.ai/lightning-ai/studios/benchmark-cloud-data-loading-libraries) | Image & Label | 10 | 1 | [Imagenet 1M](https://paperswithcode.com/sota/image-classification-on-imagenet?tag_filter=171) |
|
|
2408
2567
|
| [Optimize GeoSpatial data for model training](https://lightning.ai/lightning-ai/studios/convert-spatial-data-to-lightning-streaming) | Image & Mask | 120 | 32 | [Chesapeake Roads Spatial Context](https://github.com/isaaccorley/chesapeakersc) |
|
|
2409
|
-
| [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://
|
|
2568
|
+
| [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
|
|
2410
2569
|
| [Optimize parquet files for model training](https://lightning.ai/lightning-ai/studios/convert-parquets-to-lightning-streaming) | Parquet Files | 12 | 16 | Randomly Generated data |
|
|
2411
2570
|
|
|
2412
2571
|
|