litdata 0.2.66__tar.gz → 0.2.67__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {litdata-0.2.66/src/litdata.egg-info → litdata-0.2.67}/PKG-INFO +185 -64
- {litdata-0.2.66 → litdata-0.2.67}/README.md +183 -63
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__about__.py +1 -1
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/data_processor.py +19 -10
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/functions.py +18 -1
- litdata-0.2.67/src/litdata/raw/dataset.py +1712 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/indexer.py +63 -19
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/client.py +24 -10
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/downloader.py +46 -27
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/resolver.py +16 -2
- {litdata-0.2.66 → litdata-0.2.67/src/litdata.egg-info}/PKG-INFO +185 -64
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/requires.txt +3 -0
- litdata-0.2.66/src/litdata/raw/dataset.py +0 -223
- {litdata-0.2.66 → litdata-0.2.67}/CONTRIBUTING.md +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/LICENSE +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/MANIFEST.in +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/requirements.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/setup.cfg +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/setup.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__main__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/commands.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/cache.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/optimize.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/parser.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/constants.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/debugger.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/helpers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/imports.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/readers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/utilities.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/types.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/requirements.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/async_prefetch.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/cache.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/combined.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/compression.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/config.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/dataloader.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/dataset.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/fs_provider.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/item_loader.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/parallel.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/reader.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/sampler.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/serializers.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/shuffle.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/timing.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/writer.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/__init__.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/_pytree.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/base.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/breakpoint.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/broadcast.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/dataset_utilities.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/encryption.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/env.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/format.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/hf_dataset.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/packing.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/parquet.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/shuffle.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/subsample.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/torch_utils.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/train_test_split.py +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/SOURCES.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/dependency_links.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/entry_points.txt +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/not-zip-safe +0 -0
- {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: litdata
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.67
|
|
4
4
|
Summary: The Deep Learning framework to train, deploy, and ship AI products Lightning fast.
|
|
5
5
|
Home-page: https://github.com/Lightning-AI/litdata
|
|
6
6
|
Download-URL: https://github.com/Lightning-AI/litdata
|
|
@@ -44,6 +44,7 @@ Requires-Dist: pillow; extra == "extras"
|
|
|
44
44
|
Requires-Dist: polars; extra == "extras"
|
|
45
45
|
Requires-Dist: pyarrow<25.0.0; extra == "extras"
|
|
46
46
|
Requires-Dist: tqdm; extra == "extras"
|
|
47
|
+
Requires-Dist: uvloop; sys_platform != "win32" and extra == "extras"
|
|
47
48
|
Requires-Dist: viztracer; extra == "extras"
|
|
48
49
|
Dynamic: author
|
|
49
50
|
Dynamic: author-email
|
|
@@ -71,12 +72,13 @@ Dynamic: summary
|
|
|
71
72
|
|
|
72
73
|
|
|
73
74
|
<pre>
|
|
74
|
-
Transform Optimize
|
|
75
|
+
Transform Optimize / Stream
|
|
75
76
|
|
|
76
|
-
✅ Parallelize data processing ✅ Stream
|
|
77
|
-
✅ Create vector embeddings ✅
|
|
78
|
-
✅ Run distributed inference ✅
|
|
79
|
-
✅ Scrape websites at scale ✅
|
|
77
|
+
✅ Parallelize data processing ✅ Stream raw files with no prep
|
|
78
|
+
✅ Create vector embeddings ✅ Stream large cloud datasets
|
|
79
|
+
✅ Run distributed inference ✅ Accelerate training by 20x
|
|
80
|
+
✅ Scrape websites at scale ✅ Pause and resume data streaming
|
|
81
|
+
✅ Use remote data without local loading
|
|
80
82
|
</pre>
|
|
81
83
|
|
|
82
84
|
---
|
|
@@ -92,6 +94,7 @@ Transform Optimize
|
|
|
92
94
|
<a href="#speed-up-model-training">Optimize data</a> •
|
|
93
95
|
<a href="#transform-datasets">Transform data</a> •
|
|
94
96
|
<a href="#key-features">Features</a> •
|
|
97
|
+
<a href="#stream-raw">Stream raw files</a> •
|
|
95
98
|
<a href="#resolve-paths">Paths & cloud URLs</a> •
|
|
96
99
|
<a href="#benchmarks">Benchmarks</a> •
|
|
97
100
|
<a href="#start-from-a-template">Templates</a> •
|
|
@@ -146,6 +149,8 @@ Install all the extras
|
|
|
146
149
|
pip install 'litdata[extras]'
|
|
147
150
|
```
|
|
148
151
|
|
|
152
|
+
On Linux/macOS, `[extras]` includes optional `uvloop` for a faster asyncio event loop used by `StreamingRawDataset` (stdlib asyncio is the fallback when it is not installed).
|
|
153
|
+
|
|
149
154
|
</details>
|
|
150
155
|
|
|
151
156
|
<details>
|
|
@@ -168,29 +173,38 @@ Source: [`.claude/skills/litdata/`](.claude/skills/litdata/) in this repository
|
|
|
168
173
|
# Speed up model training
|
|
169
174
|
Stream datasets directly from cloud storage without local downloads. Choose the approach that fits your workflow:
|
|
170
175
|
|
|
171
|
-
## Option 1:
|
|
172
|
-
|
|
176
|
+
## Option 1: Stream existing files as-is ⚡⚡ — `StreamingRawDataset`
|
|
177
|
+
|
|
178
|
+
**No optimize step.** Point LitData at a folder of images, audio, text, or any files (local or cloud) and train with a normal PyTorch `DataLoader`. Downloads are **fully asynchronous** and **batched**; cloud clients include **built-in retries**. You receive **raw `bytes`** — decode, parse, or transform however you want.
|
|
179
|
+
|
|
180
|
+
Details → [Stream raw files](#stream-raw).
|
|
173
181
|
|
|
174
182
|
```python
|
|
175
183
|
from litdata import StreamingRawDataset
|
|
176
184
|
from torch.utils.data import DataLoader
|
|
185
|
+
from PIL import Image
|
|
186
|
+
import io
|
|
177
187
|
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
188
|
+
dataset = StreamingRawDataset(
|
|
189
|
+
"s3://my-bucket/raw-images/", # or gs://, azure://, /teamspace/s3_connections/..., local path
|
|
190
|
+
transform=lambda b: Image.open(io.BytesIO(b)).convert("RGB"), # optional — default is raw bytes
|
|
191
|
+
)
|
|
192
|
+
loader = DataLoader(dataset, batch_size=32, num_workers=8)
|
|
181
193
|
|
|
182
|
-
for batch in
|
|
183
|
-
|
|
184
|
-
pass
|
|
194
|
+
for batch in loader:
|
|
195
|
+
train_step(batch)
|
|
185
196
|
```
|
|
186
197
|
|
|
187
198
|
**Key benefits:**
|
|
188
199
|
|
|
189
|
-
✅ **
|
|
190
|
-
✅ **
|
|
191
|
-
✅ **
|
|
192
|
-
✅ **
|
|
193
|
-
✅ **Cloud-native:**
|
|
200
|
+
✅ **Zero preprocess:** No chunking job — use the files you already have.
|
|
201
|
+
✅ **Raw bytes, your rules:** Each sample is file `bytes`; decode with PIL, torchaudio, json, or any custom logic (`transform=` optional).
|
|
202
|
+
✅ **Fully async + batched:** Concurrent downloads via `asyncio` / `__getitems__` (not one-file-at-a-time).
|
|
203
|
+
✅ **Built-in retries:** Cloud downloads retry transient failures (adaptive client retries).
|
|
204
|
+
✅ **Cloud-native:** S3 / GCS / Azure / Studio connections; same path resolver as optimized streaming.
|
|
205
|
+
✅ **Grouped samples:** Override `setup()` to yield image+mask, audio+transcript, etc.
|
|
206
|
+
✅ **Indexed once:** `index.json.zstd` cached locally and on the bucket for fast restarts.
|
|
207
|
+
✅ **Upgrade path:** When I/O becomes the bottleneck, `optimize` → `StreamingDataset` for max throughput.
|
|
194
208
|
|
|
195
209
|
## Option 2: Optimize for maximum performance ⚡⚡⚡
|
|
196
210
|
Accelerate model training (20x faster) by optimizing datasets for streaming directly from cloud storage. Work with remote data without local downloads with features like loading data subsets, accessing individual samples, and resumable streaming.
|
|
@@ -227,7 +241,7 @@ if __name__ == "__main__":
|
|
|
227
241
|
inputs=list(range(1000)), # the inputs to the function (here it's a list of numbers)
|
|
228
242
|
output_dir="fast_data", # optimized data is stored here
|
|
229
243
|
num_workers=4, # the number of workers on the same machine
|
|
230
|
-
chunk_bytes="64MB" #
|
|
244
|
+
chunk_bytes="64MB" # default; see FAQ for larger samples
|
|
231
245
|
)
|
|
232
246
|
```
|
|
233
247
|
|
|
@@ -325,79 +339,158 @@ ld.map(
|
|
|
325
339
|
## Features for optimizing and streaming datasets for model training
|
|
326
340
|
|
|
327
341
|
<details>
|
|
328
|
-
<summary> ✅ Stream raw
|
|
342
|
+
<summary> ✅ Stream raw files as-is (no optimize) — StreamingRawDataset <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
|
|
329
343
|
|
|
330
344
|
|
|
331
|
-
|
|
345
|
+
`StreamingRawDataset` streams **your existing files** from local disk or cloud storage with **no conversion step**. It is a map-style `torch.utils.data.Dataset`: use a standard PyTorch `DataLoader` (not `StreamingDataLoader`).
|
|
332
346
|
|
|
333
|
-
**
|
|
347
|
+
**You get raw `bytes`.** LitData does not impose a sample schema — open images with PIL, parse JSONL, decode audio, run your own tokenizer, or pass a `transform=` if you prefer. Grouped items yield `list[bytes]` (e.g. image + mask).
|
|
334
348
|
|
|
335
|
-
|
|
349
|
+
Downloads are **fully asynchronous** and **batched**: when the DataLoader requests a batch, `__getitems__` fetches those files concurrently with `asyncio.gather`. Cloud clients include **built-in retries** for transient network errors.
|
|
336
350
|
|
|
337
|
-
|
|
338
|
-
# for aws s3
|
|
339
|
-
pip install "litdata[extra]" s3fs
|
|
351
|
+
Use it when you want to train or prototype on JPEGs, masks, audio, JSONL, etc. **immediately**. Switch to [`optimize` → `StreamingDataset`](#speed-up-model-training) later if you need maximum cloud training throughput.
|
|
340
352
|
|
|
341
|
-
|
|
342
|
-
|
|
353
|
+
| | `StreamingRawDataset` | `StreamingDataset` (optimized) |
|
|
354
|
+
|--|----------------------|--------------------------------|
|
|
355
|
+
| Prep | None — point at a folder | One-time `optimize` → `chunk-*.bin` + `index.json` |
|
|
356
|
+
| Item | **Raw file `bytes`** (you decide how to decode) | Deserialized samples (dict/tensor/…) |
|
|
357
|
+
| I/O | Fully async, batched downloads + retries | Chunk prefetch / cache pipeline |
|
|
358
|
+
| Loader | `torch.utils.data.DataLoader` | Prefer `StreamingDataLoader` (shuffle, resume) |
|
|
359
|
+
| Best for | Instant start, full control over bytes | Highest sustained training I/O |
|
|
360
|
+
|
|
361
|
+
### Install (cloud)
|
|
362
|
+
|
|
363
|
+
```bash
|
|
364
|
+
pip install "litdata[extra]" s3fs # Amazon S3
|
|
365
|
+
pip install "litdata[extra]" gcsfs # Google Cloud Storage
|
|
366
|
+
# Azure / Studio connections: see Paths & cloud URLs
|
|
343
367
|
```
|
|
344
368
|
|
|
345
|
-
|
|
369
|
+
### Quick start
|
|
370
|
+
|
|
346
371
|
```python
|
|
347
372
|
from torch.utils.data import DataLoader
|
|
348
373
|
from litdata import StreamingRawDataset
|
|
374
|
+
from PIL import Image
|
|
375
|
+
import io
|
|
349
376
|
|
|
350
|
-
|
|
377
|
+
def to_image(data: bytes):
|
|
378
|
+
return Image.open(io.BytesIO(data)).convert("RGB")
|
|
379
|
+
|
|
380
|
+
dataset = StreamingRawDataset(
|
|
381
|
+
"s3://my-bucket/images/", # also: gs://, azure://, /teamspace/s3_connections/..., local path
|
|
382
|
+
transform=to_image, # optional; default yields raw bytes
|
|
383
|
+
storage_options={}, # optional cloud credentials / endpoint
|
|
384
|
+
)
|
|
385
|
+
loader = DataLoader(dataset, batch_size=32, num_workers=8)
|
|
351
386
|
|
|
352
|
-
# Use with PyTorch DataLoader
|
|
353
|
-
loader = DataLoader(dataset, batch_size=32)
|
|
354
387
|
for batch in loader:
|
|
355
|
-
|
|
356
|
-
pass
|
|
388
|
+
train_step(batch)
|
|
357
389
|
```
|
|
358
390
|
|
|
359
|
-
|
|
391
|
+
### Constructor knobs
|
|
360
392
|
|
|
393
|
+
| Arg | Default | Purpose |
|
|
394
|
+
|-----|---------|---------|
|
|
395
|
+
| `input_dir` | required | Folder URL/path (same [resolver](#resolve-paths) as optimized streaming) |
|
|
396
|
+
| `cache_dir` | LitData default cache | Where the file index (and optional file cache) live |
|
|
397
|
+
| `cache_files` | `False` | If `True`, keep downloaded files on disk under `cache_dir` (mirror remote layout) |
|
|
398
|
+
| `recompute_index` | `False` | Force re-scan when remote files changed |
|
|
399
|
+
| `transform` | `None` | `fn(bytes) -> Any` or `fn(list[bytes]) -> Any` for grouped items |
|
|
400
|
+
| `storage_options` | `{}` | Cloud client options |
|
|
401
|
+
| `indexer` | `FileIndexer()` | Custom discovery (subclass `BaseIndexer`) |
|
|
402
|
+
| `max_concurrent_downloads` | `None` (adaptive) | Per-worker in-flight downloads. `None` = size-aware budget (bandwidth; Little’s-law only for medians <~8 MiB) split across workers; single-process capped at 128. An explicit `int` is used exactly (no silent clamp) |
|
|
403
|
+
| `max_prefetch` | `16` | Per-worker sequential look-ahead after each batch (default on). When `num_workers > 1`, effective look-ahead is `min(max_prefetch, 64 // num_workers)` so aggregate stays ~64 items. Pass `0` to disable |
|
|
404
|
+
| `prefetch_cache_size` | auto | LRU cap for prefetched items (defaults from `max_prefetch`) |
|
|
405
|
+
| `hedge_delay` | `0` | Seconds before a hedged duplicate GET for a slow download (`0` = off, default; opt-in) |
|
|
406
|
+
| `range_parallel_threshold` | `0` | Objects ≥ this many bytes use parallel ranged GETs (`0` = whole-object only; opt-in) |
|
|
407
|
+
| `item_type` | `"bytes"` | `"bytes"` buffers in RAM; `"path"` returns local cache paths (`cache_files=True` required) |
|
|
408
|
+
|
|
409
|
+
### Group related files (`setup`)
|
|
361
410
|
|
|
362
|
-
|
|
411
|
+
Default: **one file = one sample**. Override `setup` to filter or group (image + mask, audio + transcript, …). Return either a list of `FileMetadata` or a list of groups (`list[list[FileMetadata]]`).
|
|
363
412
|
|
|
364
413
|
```python
|
|
365
|
-
from
|
|
414
|
+
from collections import defaultdict
|
|
366
415
|
from torch.utils.data import DataLoader
|
|
367
416
|
from litdata import StreamingRawDataset
|
|
368
417
|
from litdata.raw.indexer import FileMetadata
|
|
369
418
|
|
|
370
419
|
class SegmentationRawDataset(StreamingRawDataset):
|
|
371
|
-
def setup(self, files: list[FileMetadata]) ->
|
|
372
|
-
#
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
420
|
+
def setup(self, files: list[FileMetadata]) -> list[list[FileMetadata]]:
|
|
421
|
+
# Pair img_001.jpg with img_001.png (mask) by stem
|
|
422
|
+
by_stem: dict[str, dict[str, FileMetadata]] = defaultdict(dict)
|
|
423
|
+
for f in files:
|
|
424
|
+
name = f.path.rsplit("/", 1)[-1]
|
|
425
|
+
stem, _, ext = name.rpartition(".")
|
|
426
|
+
by_stem[stem][ext.lower()] = f
|
|
427
|
+
items = []
|
|
428
|
+
for stem, parts in sorted(by_stem.items()):
|
|
429
|
+
if "jpg" in parts and "png" in parts:
|
|
430
|
+
items.append([parts["jpg"], parts["png"]])
|
|
431
|
+
return items
|
|
432
|
+
|
|
433
|
+
dataset = SegmentationRawDataset(
|
|
434
|
+
"s3://bucket/seg/",
|
|
435
|
+
transform=lambda pair: (pair[0], pair[1]), # list[bytes]: [image, mask]
|
|
436
|
+
)
|
|
437
|
+
loader = DataLoader(dataset, batch_size=16, num_workers=4)
|
|
438
|
+
for images, masks in loader:
|
|
439
|
+
...
|
|
385
440
|
```
|
|
386
441
|
|
|
387
|
-
|
|
442
|
+
### Index caching (`index.json.zstd`)
|
|
388
443
|
|
|
389
|
-
|
|
444
|
+
First open scans the tree and writes a compressed file list:
|
|
390
445
|
|
|
391
|
-
|
|
392
|
-
- **
|
|
393
|
-
- **Remote:** Automatically saved to cloud storage (e.g., `s3://bucket/files/index.json.zstd`) for reuse
|
|
446
|
+
- **Local cache** under your LitData cache dir (fast restart on the same machine)
|
|
447
|
+
- **Remote copy** next to the data when possible (e.g. `s3://bucket/files/index.json.zstd`) so every machine skips the scan
|
|
394
448
|
|
|
395
|
-
**Force Rebuild:**
|
|
396
449
|
```python
|
|
397
|
-
#
|
|
450
|
+
# After adding/removing files on the bucket:
|
|
398
451
|
dataset = StreamingRawDataset("s3://bucket/files/", recompute_index=True)
|
|
399
452
|
```
|
|
400
453
|
|
|
454
|
+
Do **not** confuse this with optimized LitData’s `index.json` (chunk metadata). Raw indexing only lists files.
|
|
455
|
+
|
|
456
|
+
### How downloads work
|
|
457
|
+
|
|
458
|
+
1. DataLoader asks for a batch of indices → `__getitems__`.
|
|
459
|
+
2. LitData **asynchronously** downloads those files **in parallel** (`asyncio.gather` + `adownload_fileobj`).
|
|
460
|
+
3. Cloud SDKs apply **retries** on transient failures (e.g. S3 adaptive retries).
|
|
461
|
+
4. Each item is returned as **`bytes`** (or `list[bytes]` if `setup` grouped files), then optional `transform`.
|
|
462
|
+
|
|
463
|
+
Your training loop stays normal PyTorch — no async/`await` in user code.
|
|
464
|
+
|
|
465
|
+
```python
|
|
466
|
+
# Default: you own the bytes
|
|
467
|
+
dataset = StreamingRawDataset("s3://bucket/files/")
|
|
468
|
+
raw: bytes = dataset[0]
|
|
469
|
+
# e.g. Image.open(io.BytesIO(raw)), json.loads(raw), np.frombuffer(raw), ...
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
### Tips
|
|
473
|
+
|
|
474
|
+
- Prefer `num_workers > 0` so worker processes overlap async batch downloads with training. Scale workers toward host vCPUs for network-bound JPEG-sized objects — avoid saturating every vCPU.
|
|
475
|
+
- On Linux, after any parent-process dataset I/O, use `DataLoader(..., multiprocessing_context="spawn", persistent_workers=True)` — default `fork` can hang S3 clients in workers.
|
|
476
|
+
- Default `max_prefetch=16` enables sequential look-ahead **per DataLoader worker**; shuffled access disables it. Pass `0` to turn off. When `num_workers > 1`, look-ahead and download concurrency both scale down with worker count so aggregate in-flight work stays bounded.
|
|
477
|
+
- Prefer an `s3://` / `gs://` URL or `/teamspace/s3_connections/...` so LitData hits the bucket directly ([resolver](#resolve-paths)) — avoid reading through FUSE.
|
|
478
|
+
- Leave `range_parallel_threshold=0` (default) for typical JPEGs; raise it only for large objects where parallel ranged GETs help.
|
|
479
|
+
- Best for medium/large files. Tiny objects (≲100 KB) are request-overhead bound — pack with [`optimize`](#speed-up-model-training) → `StreamingDataset` when I/O plateaus.
|
|
480
|
+
|
|
481
|
+
### Throughput
|
|
482
|
+
|
|
483
|
+
On ImageNet val raw over S3 (50 k JPEGs, batch size 64, spawn workers), throughput gains are clearest at **low worker counts / notebooks** (**+20–80%** at ≤8 workers). At **high workers** (≥16), results are roughly **parity within run-to-run noise**.
|
|
484
|
+
|
|
485
|
+
| workers | before | after | Δ |
|
|
486
|
+
|--------:|-------:|------:|--:|
|
|
487
|
+
| 0 | 543 | 735 | **+35%** |
|
|
488
|
+
| 2 | 816 | 1475 | **+81%** |
|
|
489
|
+
| 8 | 4841 | 5718 | **+18%** |
|
|
490
|
+
| 16+ | ~6k | ~6k | ~parity |
|
|
491
|
+
|
|
492
|
+
Useful knobs: `num_workers`, `max_prefetch` (default 16; worker-aware), `download_timeout` (batch-level hang protection). Ranged parallel downloads stay opt-in (`range_parallel_threshold=0`).
|
|
493
|
+
|
|
401
494
|
</details>
|
|
402
495
|
|
|
403
496
|
<details>
|
|
@@ -715,6 +808,35 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
|
|
|
715
808
|
|
|
716
809
|
</details>
|
|
717
810
|
|
|
811
|
+
<details>
|
|
812
|
+
<summary> ✅ FAQ: chunk size & shuffle before optimize <a id="faq-chunk-shuffle" href="#faq-chunk-shuffle">🔗</a> </summary>
|
|
813
|
+
|
|
814
|
+
|
|
815
|
+
### What `chunk_bytes` should I use?
|
|
816
|
+
|
|
817
|
+
Default is **64MB** — a good starting point for typical small/medium samples.
|
|
818
|
+
|
|
819
|
+
When each datapoint is large (e.g. a few MB), prefer a **larger chunk** (practical range often **256–512MB**) so each chunk holds more samples and **intra-chunk batch randomization** has a bigger pool. Tradeoff: larger chunks take **longer to download** before they can be used.
|
|
820
|
+
|
|
821
|
+
This is expert guidance (recommended-range mindset), not a published chunk-size sweep.
|
|
822
|
+
|
|
823
|
+
### Is StreamingDataset shuffle enough if my source data is ordered?
|
|
824
|
+
|
|
825
|
+
**Not always.** LitData handles **distributed sampling** and **bucket sampling within chunks** automatically (`shuffle=True` randomizes chunk order and item order inside each chunk). That is **not** a substitute for a fully shuffled file-level DataLoader when the source has strong structure (same subject/set contiguous, class blocks, etc.).
|
|
826
|
+
|
|
827
|
+
If ordered data would make chunked sampling problematic and you cannot embed the grouping as the sample unit:
|
|
828
|
+
|
|
829
|
+
- Shuffle the list of samples **before** `optimize` so chunks mix well, **or**
|
|
830
|
+
- Use [`StreamingRawDataset`](#stream-raw) (per-file random access via a standard PyTorch `DataLoader` with `shuffle=True`) instead of optimize → `StreamingDataset`.
|
|
831
|
+
|
|
832
|
+
### FUSE vs LitData (Lightning Studios)
|
|
833
|
+
|
|
834
|
+
`/teamspace/s3_connections` (and related mounts) are **FUSE** — fine for browsing, not for training I/O. Under load they are very slow and can crash. Pass the same path into LitData (`StreamingRawDataset` / `StreamingDataset` / `optimize`): LitData resolves it and talks **directly** to the bucket ([Resolve any path](#resolve-paths)).
|
|
835
|
+
|
|
836
|
+
Rough ImageNet order-of-magnitude on a Studio (not hard guarantees; right tuning for raw): FUSE hand-read ~**600** images/s · [`StreamingRawDataset`](#stream-raw) ~**6–7k** · optimized [`StreamingDataset`](#speed-up-model-training) (64MB chunks) ~**11k**.
|
|
837
|
+
|
|
838
|
+
</details>
|
|
839
|
+
|
|
718
840
|
<details>
|
|
719
841
|
<summary> ✅ StreamingDataset & StreamingDataLoader knobs <a id="streaming-kwargs" href="#streaming-kwargs">🔗</a> </summary>
|
|
720
842
|
|
|
@@ -1093,7 +1215,7 @@ if __name__ == "__main__":
|
|
|
1093
1215
|
|
|
1094
1216
|
Mix and match different sets of data to experiment and create better models.
|
|
1095
1217
|
|
|
1096
|
-
Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://
|
|
1218
|
+
Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
|
|
1097
1219
|
|
|
1098
1220
|
```python
|
|
1099
1221
|
from litdata import StreamingDataset, CombinedStreamingDataset, StreamingDataLoader, TokensLoader
|
|
@@ -2213,7 +2335,7 @@ Full knob list for `litdata.optimize` (see Quick start for the minimal recipe).
|
|
|
2213
2335
|
| `output_dir` | `"optimized_data"` | Local or cloud ([resolver](#resolve-paths)); version remote prefixes |
|
|
2214
2336
|
| `input_dir` | `None` | Remote input root for background download |
|
|
2215
2337
|
| `weights` | `None` | Per-input weights to balance workers |
|
|
2216
|
-
| `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"
|
|
2338
|
+
| `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`; see [FAQ](#faq-chunk-shuffle) for larger samples) |
|
|
2217
2339
|
| `chunk_size` | `None` | Max items (or tokens with `TokensLoader`) per chunk |
|
|
2218
2340
|
| `align_chunking` | `False` | Match single-worker chunk boundaries (needs `chunk_size`; uneven load) |
|
|
2219
2341
|
| `compression` | `None` | `"zstd"` today |
|
|
@@ -2316,8 +2438,7 @@ Speed to stream raw Imagenet 1.2M from different cloud storage providers:
|
|
|
2316
2438
|
| AWS S3 | ~6400 +/- 100 | ~3200 +/- 100 |
|
|
2317
2439
|
| Google Cloud Storage | ~5650 +/- 100 | ~3100 +/- 100 |
|
|
2318
2440
|
|
|
2319
|
-
> **
|
|
2320
|
-
> Use `StreamingRawDataset` if you want to stream your data as-is. Use `StreamingDataset` if you want the fastest streaming and are okay with optimizing your data first.
|
|
2441
|
+
> **Also see:** [`StreamingRawDataset`](#stream-raw) streams existing files with **no optimize step** (great default to start). Use `StreamingDataset` after `optimize` when you need the highest sustained training throughput.
|
|
2321
2442
|
|
|
2322
2443
|
|
|
2323
2444
|
|
|
@@ -2406,7 +2527,7 @@ Below are templates for real-world applications of LitData at scale.
|
|
|
2406
2527
|
| -------------------------------- | ----------------- | ----------------- | -------------- | -------------- |
|
|
2407
2528
|
| [Benchmark cloud data-loading libraries](https://lightning.ai/lightning-ai/studios/benchmark-cloud-data-loading-libraries) | Image & Label | 10 | 1 | [Imagenet 1M](https://paperswithcode.com/sota/image-classification-on-imagenet?tag_filter=171) |
|
|
2408
2529
|
| [Optimize GeoSpatial data for model training](https://lightning.ai/lightning-ai/studios/convert-spatial-data-to-lightning-streaming) | Image & Mask | 120 | 32 | [Chesapeake Roads Spatial Context](https://github.com/isaaccorley/chesapeakersc) |
|
|
2409
|
-
| [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://
|
|
2530
|
+
| [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
|
|
2410
2531
|
| [Optimize parquet files for model training](https://lightning.ai/lightning-ai/studios/convert-parquets-to-lightning-streaming) | Parquet Files | 12 | 16 | Randomly Generated data |
|
|
2411
2532
|
|
|
2412
2533
|
|