litdata 0.2.66__tar.gz → 0.2.67__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (75) hide show
  1. {litdata-0.2.66/src/litdata.egg-info → litdata-0.2.67}/PKG-INFO +185 -64
  2. {litdata-0.2.66 → litdata-0.2.67}/README.md +183 -63
  3. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__about__.py +1 -1
  4. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/data_processor.py +19 -10
  5. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/functions.py +18 -1
  6. litdata-0.2.67/src/litdata/raw/dataset.py +1712 -0
  7. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/indexer.py +63 -19
  8. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/client.py +24 -10
  9. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/downloader.py +46 -27
  10. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/resolver.py +16 -2
  11. {litdata-0.2.66 → litdata-0.2.67/src/litdata.egg-info}/PKG-INFO +185 -64
  12. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/requires.txt +3 -0
  13. litdata-0.2.66/src/litdata/raw/dataset.py +0 -223
  14. {litdata-0.2.66 → litdata-0.2.67}/CONTRIBUTING.md +0 -0
  15. {litdata-0.2.66 → litdata-0.2.67}/LICENSE +0 -0
  16. {litdata-0.2.66 → litdata-0.2.67}/MANIFEST.in +0 -0
  17. {litdata-0.2.66 → litdata-0.2.67}/requirements.txt +0 -0
  18. {litdata-0.2.66 → litdata-0.2.67}/setup.cfg +0 -0
  19. {litdata-0.2.66 → litdata-0.2.67}/setup.py +0 -0
  20. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__init__.py +0 -0
  21. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/__main__.py +0 -0
  22. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/__init__.py +0 -0
  23. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/commands.py +0 -0
  24. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/__init__.py +0 -0
  25. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/cache.py +0 -0
  26. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/handler/optimize.py +0 -0
  27. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/cli/parser.py +0 -0
  28. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/constants.py +0 -0
  29. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/debugger.py +0 -0
  30. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/helpers.py +0 -0
  31. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/imports.py +0 -0
  32. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/__init__.py +0 -0
  33. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/readers.py +0 -0
  34. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/processing/utilities.py +0 -0
  35. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/__init__.py +0 -0
  36. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/raw/types.py +0 -0
  37. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/requirements.py +0 -0
  38. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/__init__.py +0 -0
  39. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/async_prefetch.py +0 -0
  40. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/cache.py +0 -0
  41. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/combined.py +0 -0
  42. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/compression.py +0 -0
  43. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/config.py +0 -0
  44. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/dataloader.py +0 -0
  45. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/dataset.py +0 -0
  46. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/fs_provider.py +0 -0
  47. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/item_loader.py +0 -0
  48. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/parallel.py +0 -0
  49. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/reader.py +0 -0
  50. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/sampler.py +0 -0
  51. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/serializers.py +0 -0
  52. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/shuffle.py +0 -0
  53. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/timing.py +0 -0
  54. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/streaming/writer.py +0 -0
  55. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/__init__.py +0 -0
  56. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/_pytree.py +0 -0
  57. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/base.py +0 -0
  58. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/breakpoint.py +0 -0
  59. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/broadcast.py +0 -0
  60. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/dataset_utilities.py +0 -0
  61. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/encryption.py +0 -0
  62. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/env.py +0 -0
  63. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/format.py +0 -0
  64. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/hf_dataset.py +0 -0
  65. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/packing.py +0 -0
  66. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/parquet.py +0 -0
  67. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/shuffle.py +0 -0
  68. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/subsample.py +0 -0
  69. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/torch_utils.py +0 -0
  70. {litdata-0.2.66 → litdata-0.2.67}/src/litdata/utilities/train_test_split.py +0 -0
  71. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/SOURCES.txt +0 -0
  72. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/dependency_links.txt +0 -0
  73. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/entry_points.txt +0 -0
  74. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/not-zip-safe +0 -0
  75. {litdata-0.2.66 → litdata-0.2.67}/src/litdata.egg-info/top_level.txt +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: litdata
3
- Version: 0.2.66
3
+ Version: 0.2.67
4
4
  Summary: The Deep Learning framework to train, deploy, and ship AI products Lightning fast.
5
5
  Home-page: https://github.com/Lightning-AI/litdata
6
6
  Download-URL: https://github.com/Lightning-AI/litdata
@@ -44,6 +44,7 @@ Requires-Dist: pillow; extra == "extras"
44
44
  Requires-Dist: polars; extra == "extras"
45
45
  Requires-Dist: pyarrow<25.0.0; extra == "extras"
46
46
  Requires-Dist: tqdm; extra == "extras"
47
+ Requires-Dist: uvloop; sys_platform != "win32" and extra == "extras"
47
48
  Requires-Dist: viztracer; extra == "extras"
48
49
  Dynamic: author
49
50
  Dynamic: author-email
@@ -71,12 +72,13 @@ Dynamic: summary
71
72
  &nbsp;
72
73
 
73
74
  <pre>
74
- Transform Optimize
75
+ Transform Optimize / Stream
75
76
 
76
- ✅ Parallelize data processing ✅ Stream large cloud datasets
77
- ✅ Create vector embeddings ✅ Accelerate training by 20x
78
- ✅ Run distributed inference ✅ Pause and resume data streaming
79
- ✅ Scrape websites at scale ✅ Use remote data without local loading
77
+ ✅ Parallelize data processing ✅ Stream raw files with no prep
78
+ ✅ Create vector embeddings ✅ Stream large cloud datasets
79
+ ✅ Run distributed inference ✅ Accelerate training by 20x
80
+ ✅ Scrape websites at scale ✅ Pause and resume data streaming
81
+ ✅ Use remote data without local loading
80
82
  </pre>
81
83
 
82
84
  ---
@@ -92,6 +94,7 @@ Transform Optimize
92
94
  <a href="#speed-up-model-training">Optimize data</a> •
93
95
  <a href="#transform-datasets">Transform data</a> •
94
96
  <a href="#key-features">Features</a> •
97
+ <a href="#stream-raw">Stream raw files</a> •
95
98
  <a href="#resolve-paths">Paths & cloud URLs</a> •
96
99
  <a href="#benchmarks">Benchmarks</a> •
97
100
  <a href="#start-from-a-template">Templates</a> •
@@ -146,6 +149,8 @@ Install all the extras
146
149
  pip install 'litdata[extras]'
147
150
  ```
148
151
 
152
+ On Linux/macOS, `[extras]` includes optional `uvloop` for a faster asyncio event loop used by `StreamingRawDataset` (stdlib asyncio is the fallback when it is not installed).
153
+
149
154
  </details>
150
155
 
151
156
  <details>
@@ -168,29 +173,38 @@ Source: [`.claude/skills/litdata/`](.claude/skills/litdata/) in this repository
168
173
  # Speed up model training
169
174
  Stream datasets directly from cloud storage without local downloads. Choose the approach that fits your workflow:
170
175
 
171
- ## Option 1: Start immediately with existing data ⚡⚡
172
- Stream raw files directly from cloud storage - no pre-optimization needed.
176
+ ## Option 1: Stream existing files as-is ⚡⚡ — `StreamingRawDataset`
177
+
178
+ **No optimize step.** Point LitData at a folder of images, audio, text, or any files (local or cloud) and train with a normal PyTorch `DataLoader`. Downloads are **fully asynchronous** and **batched**; cloud clients include **built-in retries**. You receive **raw `bytes`** — decode, parse, or transform however you want.
179
+
180
+ Details → [Stream raw files](#stream-raw).
173
181
 
174
182
  ```python
175
183
  from litdata import StreamingRawDataset
176
184
  from torch.utils.data import DataLoader
185
+ from PIL import Image
186
+ import io
177
187
 
178
- # Point to your existing cloud data
179
- dataset = StreamingRawDataset("s3://my-bucket/raw-data/")
180
- dataloader = DataLoader(dataset, batch_size=32)
188
+ dataset = StreamingRawDataset(
189
+ "s3://my-bucket/raw-images/", # or gs://, azure://, /teamspace/s3_connections/..., local path
190
+ transform=lambda b: Image.open(io.BytesIO(b)).convert("RGB"), # optional — default is raw bytes
191
+ )
192
+ loader = DataLoader(dataset, batch_size=32, num_workers=8)
181
193
 
182
- for batch in dataloader:
183
- # Process raw bytes on-the-fly
184
- pass
194
+ for batch in loader:
195
+ train_step(batch)
185
196
  ```
186
197
 
187
198
  **Key benefits:**
188
199
 
189
- ✅ **Instant access:** Start streaming immediately without preprocessing.
190
- ✅ **Zero setup time:** No data conversion or optimization required.
191
- ✅ **Native format:** Work with original file formats (images, text, etc.).
192
- ✅ **Flexible processing:** Apply transformations on-the-fly during streaming.
193
- ✅ **Cloud-native:** Stream directly from S3, GCS, or Azure storage.
200
+ ✅ **Zero preprocess:** No chunking job — use the files you already have.
201
+ ✅ **Raw bytes, your rules:** Each sample is file `bytes`; decode with PIL, torchaudio, json, or any custom logic (`transform=` optional).
202
+ ✅ **Fully async + batched:** Concurrent downloads via `asyncio` / `__getitems__` (not one-file-at-a-time).
203
+ ✅ **Built-in retries:** Cloud downloads retry transient failures (adaptive client retries).
204
+ ✅ **Cloud-native:** S3 / GCS / Azure / Studio connections; same path resolver as optimized streaming.
205
+ ✅ **Grouped samples:** Override `setup()` to yield image+mask, audio+transcript, etc.
206
+ ✅ **Indexed once:** `index.json.zstd` cached locally and on the bucket for fast restarts.
207
+ ✅ **Upgrade path:** When I/O becomes the bottleneck, `optimize` → `StreamingDataset` for max throughput.
194
208
 
195
209
  ## Option 2: Optimize for maximum performance ⚡⚡⚡
196
210
  Accelerate model training (20x faster) by optimizing datasets for streaming directly from cloud storage. Work with remote data without local downloads with features like loading data subsets, accessing individual samples, and resumable streaming.
@@ -227,7 +241,7 @@ if __name__ == "__main__":
227
241
  inputs=list(range(1000)), # the inputs to the function (here it's a list of numbers)
228
242
  output_dir="fast_data", # optimized data is stored here
229
243
  num_workers=4, # the number of workers on the same machine
230
- chunk_bytes="64MB" # size of each chunk
244
+ chunk_bytes="64MB" # default; see FAQ for larger samples
231
245
  )
232
246
  ```
233
247
 
@@ -325,79 +339,158 @@ ld.map(
325
339
  ## Features for optimizing and streaming datasets for model training
326
340
 
327
341
  <details>
328
- <summary> ✅ Stream raw datasets from cloud storage (beta) <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
342
+ <summary> ✅ Stream raw files as-is (no optimize) — StreamingRawDataset <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
329
343
  &nbsp;
330
344
 
331
- Effortlessly stream raw files (images, text, etc.) directly from S3, GCS, and Azure cloud storage without any optimization or conversion. Ideal for workflows requiring instant access to original data in its native format.
345
+ `StreamingRawDataset` streams **your existing files** from local disk or cloud storage with **no conversion step**. It is a map-style `torch.utils.data.Dataset`: use a standard PyTorch `DataLoader` (not `StreamingDataLoader`).
332
346
 
333
- **Prerequisites:**
347
+ **You get raw `bytes`.** LitData does not impose a sample schema — open images with PIL, parse JSONL, decode audio, run your own tokenizer, or pass a `transform=` if you prefer. Grouped items yield `list[bytes]` (e.g. image + mask).
334
348
 
335
- Install the required dependencies to stream raw datasets from cloud storage like **Amazon S3** or **Google Cloud Storage**:
349
+ Downloads are **fully asynchronous** and **batched**: when the DataLoader requests a batch, `__getitems__` fetches those files concurrently with `asyncio.gather`. Cloud clients include **built-in retries** for transient network errors.
336
350
 
337
- ```bash
338
- # for aws s3
339
- pip install "litdata[extra]" s3fs
351
+ Use it when you want to train or prototype on JPEGs, masks, audio, JSONL, etc. **immediately**. Switch to [`optimize` → `StreamingDataset`](#speed-up-model-training) later if you need maximum cloud training throughput.
340
352
 
341
- # for gcloud storage
342
- pip install "litdata[extra]" gcsfs
353
+ | | `StreamingRawDataset` | `StreamingDataset` (optimized) |
354
+ |--|----------------------|--------------------------------|
355
+ | Prep | None — point at a folder | One-time `optimize` → `chunk-*.bin` + `index.json` |
356
+ | Item | **Raw file `bytes`** (you decide how to decode) | Deserialized samples (dict/tensor/…) |
357
+ | I/O | Fully async, batched downloads + retries | Chunk prefetch / cache pipeline |
358
+ | Loader | `torch.utils.data.DataLoader` | Prefer `StreamingDataLoader` (shuffle, resume) |
359
+ | Best for | Instant start, full control over bytes | Highest sustained training I/O |
360
+
361
+ ### Install (cloud)
362
+
363
+ ```bash
364
+ pip install "litdata[extra]" s3fs # Amazon S3
365
+ pip install "litdata[extra]" gcsfs # Google Cloud Storage
366
+ # Azure / Studio connections: see Paths & cloud URLs
343
367
  ```
344
368
 
345
- **Usage Example:**
369
+ ### Quick start
370
+
346
371
  ```python
347
372
  from torch.utils.data import DataLoader
348
373
  from litdata import StreamingRawDataset
374
+ from PIL import Image
375
+ import io
349
376
 
350
- dataset = StreamingRawDataset("s3://bucket/files/")
377
+ def to_image(data: bytes):
378
+ return Image.open(io.BytesIO(data)).convert("RGB")
379
+
380
+ dataset = StreamingRawDataset(
381
+ "s3://my-bucket/images/", # also: gs://, azure://, /teamspace/s3_connections/..., local path
382
+ transform=to_image, # optional; default yields raw bytes
383
+ storage_options={}, # optional cloud credentials / endpoint
384
+ )
385
+ loader = DataLoader(dataset, batch_size=32, num_workers=8)
351
386
 
352
- # Use with PyTorch DataLoader
353
- loader = DataLoader(dataset, batch_size=32)
354
387
  for batch in loader:
355
- # Each item is raw bytes
356
- pass
388
+ train_step(batch)
357
389
  ```
358
390
 
359
- > Use `StreamingRawDataset` to stream your data as-is. Use `StreamingDataset` for fastest streaming after optimizing your data.
391
+ ### Constructor knobs
360
392
 
393
+ | Arg | Default | Purpose |
394
+ |-----|---------|---------|
395
+ | `input_dir` | required | Folder URL/path (same [resolver](#resolve-paths) as optimized streaming) |
396
+ | `cache_dir` | LitData default cache | Where the file index (and optional file cache) live |
397
+ | `cache_files` | `False` | If `True`, keep downloaded files on disk under `cache_dir` (mirror remote layout) |
398
+ | `recompute_index` | `False` | Force re-scan when remote files changed |
399
+ | `transform` | `None` | `fn(bytes) -> Any` or `fn(list[bytes]) -> Any` for grouped items |
400
+ | `storage_options` | `{}` | Cloud client options |
401
+ | `indexer` | `FileIndexer()` | Custom discovery (subclass `BaseIndexer`) |
402
+ | `max_concurrent_downloads` | `None` (adaptive) | Per-worker in-flight downloads. `None` = size-aware budget (bandwidth; Little’s-law only for medians &lt;~8 MiB) split across workers; single-process capped at 128. An explicit `int` is used exactly (no silent clamp) |
403
+ | `max_prefetch` | `16` | Per-worker sequential look-ahead after each batch (default on). When `num_workers > 1`, effective look-ahead is `min(max_prefetch, 64 // num_workers)` so aggregate stays ~64 items. Pass `0` to disable |
404
+ | `prefetch_cache_size` | auto | LRU cap for prefetched items (defaults from `max_prefetch`) |
405
+ | `hedge_delay` | `0` | Seconds before a hedged duplicate GET for a slow download (`0` = off, default; opt-in) |
406
+ | `range_parallel_threshold` | `0` | Objects ≥ this many bytes use parallel ranged GETs (`0` = whole-object only; opt-in) |
407
+ | `item_type` | `"bytes"` | `"bytes"` buffers in RAM; `"path"` returns local cache paths (`cache_files=True` required) |
408
+
409
+ ### Group related files (`setup`)
361
410
 
362
- You can also customize how files are grouped by subclassing `StreamingRawDataset` and overriding the `setup` method. This is useful for pairing related files (e.g., image and mask, audio and transcript) or any custom grouping logic.
411
+ Default: **one file = one sample**. Override `setup` to filter or group (image + mask, audio + transcript, …). Return either a list of `FileMetadata` or a list of groups (`list[list[FileMetadata]]`).
363
412
 
364
413
  ```python
365
- from typing import Union
414
+ from collections import defaultdict
366
415
  from torch.utils.data import DataLoader
367
416
  from litdata import StreamingRawDataset
368
417
  from litdata.raw.indexer import FileMetadata
369
418
 
370
419
  class SegmentationRawDataset(StreamingRawDataset):
371
- def setup(self, files: list[FileMetadata]) -> Union[list[FileMetadata], list[list[FileMetadata]]]:
372
- # TODO: Implement your custom grouping logic here.
373
- # For example, group files by prefix, extension, or any rule you need.
374
- # Return a list of groups, where each group is a list of FileMetadata.
375
- # Example:
376
- # return [[image, mask], ...]
377
- pass
378
-
379
- # Initialize the custom dataset
380
- dataset = SegmentationRawDataset("s3://bucket/files/")
381
- loader = DataLoader(dataset, batch_size=32)
382
- for item in loader:
383
- # Each item in the batch is a pair: [image_bytes, mask_bytes]
384
- pass
420
+ def setup(self, files: list[FileMetadata]) -> list[list[FileMetadata]]:
421
+ # Pair img_001.jpg with img_001.png (mask) by stem
422
+ by_stem: dict[str, dict[str, FileMetadata]] = defaultdict(dict)
423
+ for f in files:
424
+ name = f.path.rsplit("/", 1)[-1]
425
+ stem, _, ext = name.rpartition(".")
426
+ by_stem[stem][ext.lower()] = f
427
+ items = []
428
+ for stem, parts in sorted(by_stem.items()):
429
+ if "jpg" in parts and "png" in parts:
430
+ items.append([parts["jpg"], parts["png"]])
431
+ return items
432
+
433
+ dataset = SegmentationRawDataset(
434
+ "s3://bucket/seg/",
435
+ transform=lambda pair: (pair[0], pair[1]), # list[bytes]: [image, mask]
436
+ )
437
+ loader = DataLoader(dataset, batch_size=16, num_workers=4)
438
+ for images, masks in loader:
439
+ ...
385
440
  ```
386
441
 
387
- **Smart Index Caching**
442
+ ### Index caching (`index.json.zstd`)
388
443
 
389
- `StreamingRawDataset` automatically caches the file index for instant startup. Initial scan, builds and caches the index, then subsequent runs load instantly.
444
+ First open scans the tree and writes a compressed file list:
390
445
 
391
- **Two-Level Cache:**
392
- - **Local:** Stored in your cache directory for instant access
393
- - **Remote:** Automatically saved to cloud storage (e.g., `s3://bucket/files/index.json.zstd`) for reuse
446
+ - **Local cache** under your LitData cache dir (fast restart on the same machine)
447
+ - **Remote copy** next to the data when possible (e.g. `s3://bucket/files/index.json.zstd`) so every machine skips the scan
394
448
 
395
- **Force Rebuild:**
396
449
  ```python
397
- # When dataset files have changed
450
+ # After adding/removing files on the bucket:
398
451
  dataset = StreamingRawDataset("s3://bucket/files/", recompute_index=True)
399
452
  ```
400
453
 
454
+ Do **not** confuse this with optimized LitData’s `index.json` (chunk metadata). Raw indexing only lists files.
455
+
456
+ ### How downloads work
457
+
458
+ 1. DataLoader asks for a batch of indices → `__getitems__`.
459
+ 2. LitData **asynchronously** downloads those files **in parallel** (`asyncio.gather` + `adownload_fileobj`).
460
+ 3. Cloud SDKs apply **retries** on transient failures (e.g. S3 adaptive retries).
461
+ 4. Each item is returned as **`bytes`** (or `list[bytes]` if `setup` grouped files), then optional `transform`.
462
+
463
+ Your training loop stays normal PyTorch — no async/`await` in user code.
464
+
465
+ ```python
466
+ # Default: you own the bytes
467
+ dataset = StreamingRawDataset("s3://bucket/files/")
468
+ raw: bytes = dataset[0]
469
+ # e.g. Image.open(io.BytesIO(raw)), json.loads(raw), np.frombuffer(raw), ...
470
+ ```
471
+
472
+ ### Tips
473
+
474
+ - Prefer `num_workers > 0` so worker processes overlap async batch downloads with training. Scale workers toward host vCPUs for network-bound JPEG-sized objects — avoid saturating every vCPU.
475
+ - On Linux, after any parent-process dataset I/O, use `DataLoader(..., multiprocessing_context="spawn", persistent_workers=True)` — default `fork` can hang S3 clients in workers.
476
+ - Default `max_prefetch=16` enables sequential look-ahead **per DataLoader worker**; shuffled access disables it. Pass `0` to turn off. When `num_workers > 1`, look-ahead and download concurrency both scale down with worker count so aggregate in-flight work stays bounded.
477
+ - Prefer an `s3://` / `gs://` URL or `/teamspace/s3_connections/...` so LitData hits the bucket directly ([resolver](#resolve-paths)) — avoid reading through FUSE.
478
+ - Leave `range_parallel_threshold=0` (default) for typical JPEGs; raise it only for large objects where parallel ranged GETs help.
479
+ - Best for medium/large files. Tiny objects (≲100 KB) are request-overhead bound — pack with [`optimize`](#speed-up-model-training) → `StreamingDataset` when I/O plateaus.
480
+
481
+ ### Throughput
482
+
483
+ On ImageNet val raw over S3 (50 k JPEGs, batch size 64, spawn workers), throughput gains are clearest at **low worker counts / notebooks** (**+20–80%** at ≤8 workers). At **high workers** (≥16), results are roughly **parity within run-to-run noise**.
484
+
485
+ | workers | before | after | Δ |
486
+ |--------:|-------:|------:|--:|
487
+ | 0 | 543 | 735 | **+35%** |
488
+ | 2 | 816 | 1475 | **+81%** |
489
+ | 8 | 4841 | 5718 | **+18%** |
490
+ | 16+ | ~6k | ~6k | ~parity |
491
+
492
+ Useful knobs: `num_workers`, `max_prefetch` (default 16; worker-aware), `download_timeout` (batch-level hang protection). Ranged parallel downloads stay opt-in (`range_parallel_threshold=0`).
493
+
401
494
  </details>
402
495
 
403
496
  <details>
@@ -715,6 +808,35 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
715
808
 
716
809
  </details>
717
810
 
811
+ <details>
812
+ <summary> ✅ FAQ: chunk size &amp; shuffle before optimize <a id="faq-chunk-shuffle" href="#faq-chunk-shuffle">🔗</a> </summary>
813
+ &nbsp;
814
+
815
+ ### What `chunk_bytes` should I use?
816
+
817
+ Default is **64MB** — a good starting point for typical small/medium samples.
818
+
819
+ When each datapoint is large (e.g. a few MB), prefer a **larger chunk** (practical range often **256–512MB**) so each chunk holds more samples and **intra-chunk batch randomization** has a bigger pool. Tradeoff: larger chunks take **longer to download** before they can be used.
820
+
821
+ This is expert guidance (recommended-range mindset), not a published chunk-size sweep.
822
+
823
+ ### Is StreamingDataset shuffle enough if my source data is ordered?
824
+
825
+ **Not always.** LitData handles **distributed sampling** and **bucket sampling within chunks** automatically (`shuffle=True` randomizes chunk order and item order inside each chunk). That is **not** a substitute for a fully shuffled file-level DataLoader when the source has strong structure (same subject/set contiguous, class blocks, etc.).
826
+
827
+ If ordered data would make chunked sampling problematic and you cannot embed the grouping as the sample unit:
828
+
829
+ - Shuffle the list of samples **before** `optimize` so chunks mix well, **or**
830
+ - Use [`StreamingRawDataset`](#stream-raw) (per-file random access via a standard PyTorch `DataLoader` with `shuffle=True`) instead of optimize → `StreamingDataset`.
831
+
832
+ ### FUSE vs LitData (Lightning Studios)
833
+
834
+ `/teamspace/s3_connections` (and related mounts) are **FUSE** — fine for browsing, not for training I/O. Under load they are very slow and can crash. Pass the same path into LitData (`StreamingRawDataset` / `StreamingDataset` / `optimize`): LitData resolves it and talks **directly** to the bucket ([Resolve any path](#resolve-paths)).
835
+
836
+ Rough ImageNet order-of-magnitude on a Studio (not hard guarantees; right tuning for raw): FUSE hand-read ~**600** images/s · [`StreamingRawDataset`](#stream-raw) ~**6–7k** · optimized [`StreamingDataset`](#speed-up-model-training) (64MB chunks) ~**11k**.
837
+
838
+ </details>
839
+
718
840
  <details>
719
841
  <summary> ✅ StreamingDataset & StreamingDataLoader knobs <a id="streaming-kwargs" href="#streaming-kwargs">🔗</a> </summary>
720
842
  &nbsp;
@@ -1093,7 +1215,7 @@ if __name__ == "__main__":
1093
1215
 
1094
1216
  Mix and match different sets of data to experiment and create better models.
1095
1217
 
1096
- Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://huggingface.co/datasets/cerebras/SlimPajama-627B) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
1218
+ Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
1097
1219
 
1098
1220
  ```python
1099
1221
  from litdata import StreamingDataset, CombinedStreamingDataset, StreamingDataLoader, TokensLoader
@@ -2213,7 +2335,7 @@ Full knob list for `litdata.optimize` (see Quick start for the minimal recipe).
2213
2335
  | `output_dir` | `"optimized_data"` | Local or cloud ([resolver](#resolve-paths)); version remote prefixes |
2214
2336
  | `input_dir` | `None` | Remote input root for background download |
2215
2337
  | `weights` | `None` | Per-input weights to balance workers |
2216
- | `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`) |
2338
+ | `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`; see [FAQ](#faq-chunk-shuffle) for larger samples) |
2217
2339
  | `chunk_size` | `None` | Max items (or tokens with `TokensLoader`) per chunk |
2218
2340
  | `align_chunking` | `False` | Match single-worker chunk boundaries (needs `chunk_size`; uneven load) |
2219
2341
  | `compression` | `None` | `"zstd"` today |
@@ -2316,8 +2438,7 @@ Speed to stream raw Imagenet 1.2M from different cloud storage providers:
2316
2438
  | AWS S3 | ~6400 +/- 100 | ~3200 +/- 100 |
2317
2439
  | Google Cloud Storage | ~5650 +/- 100 | ~3100 +/- 100 |
2318
2440
 
2319
- > **Note:**
2320
- > Use `StreamingRawDataset` if you want to stream your data as-is. Use `StreamingDataset` if you want the fastest streaming and are okay with optimizing your data first.
2441
+ > **Also see:** [`StreamingRawDataset`](#stream-raw) streams existing files with **no optimize step** (great default to start). Use `StreamingDataset` after `optimize` when you need the highest sustained training throughput.
2321
2442
 
2322
2443
  &nbsp;
2323
2444
 
@@ -2406,7 +2527,7 @@ Below are templates for real-world applications of LitData at scale.
2406
2527
  | -------------------------------- | ----------------- | ----------------- | -------------- | -------------- |
2407
2528
  | [Benchmark cloud data-loading libraries](https://lightning.ai/lightning-ai/studios/benchmark-cloud-data-loading-libraries) | Image & Label | 10 | 1 | [Imagenet 1M](https://paperswithcode.com/sota/image-classification-on-imagenet?tag_filter=171) |
2408
2529
  | [Optimize GeoSpatial data for model training](https://lightning.ai/lightning-ai/studios/convert-spatial-data-to-lightning-streaming) | Image & Mask | 120 | 32 | [Chesapeake Roads Spatial Context](https://github.com/isaaccorley/chesapeakersc) |
2409
- | [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://huggingface.co/datasets/cerebras/SlimPajama-627B) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
2530
+ | [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
2410
2531
  | [Optimize parquet files for model training](https://lightning.ai/lightning-ai/studios/convert-parquets-to-lightning-streaming) | Parquet Files | 12 | 16 | Randomly Generated data |
2411
2532
 
2412
2533
  &nbsp;