litdata 0.2.66__tar.gz → 0.2.68__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (80) hide show
  1. {litdata-0.2.66 → litdata-0.2.68}/CONTRIBUTING.md +1 -1
  2. {litdata-0.2.66/src/litdata.egg-info → litdata-0.2.68}/PKG-INFO +259 -100
  3. {litdata-0.2.66 → litdata-0.2.68}/README.md +257 -99
  4. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__about__.py +1 -1
  5. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__init__.py +6 -0
  6. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/constants.py +6 -0
  7. litdata-0.2.68/src/litdata/debugger.py +397 -0
  8. litdata-0.2.68/src/litdata/exceptions.py +36 -0
  9. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/data_processor.py +226 -26
  10. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/functions.py +53 -6
  11. litdata-0.2.68/src/litdata/raw/dataset.py +1712 -0
  12. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/indexer.py +63 -19
  13. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/async_prefetch.py +10 -2
  14. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/cache.py +2 -2
  15. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/client.py +35 -10
  16. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/combined.py +3 -12
  17. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/compression.py +28 -8
  18. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/config.py +9 -17
  19. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/dataloader.py +24 -3
  20. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/dataset.py +108 -9
  21. litdata-0.2.68/src/litdata/streaming/dataset_update.py +299 -0
  22. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/downloader.py +276 -132
  23. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/item_loader.py +317 -118
  24. litdata-0.2.68/src/litdata/streaming/posix_fast.py +396 -0
  25. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/reader.py +117 -36
  26. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/resolver.py +16 -2
  27. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/shuffle.py +52 -0
  28. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/writer.py +26 -13
  29. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/format.py +1 -1
  30. litdata-0.2.68/src/litdata/utilities/keys_index.py +868 -0
  31. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/shuffle.py +114 -0
  32. {litdata-0.2.66 → litdata-0.2.68/src/litdata.egg-info}/PKG-INFO +259 -100
  33. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/SOURCES.txt +4 -0
  34. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/requires.txt +3 -0
  35. litdata-0.2.66/src/litdata/debugger.py +0 -205
  36. litdata-0.2.66/src/litdata/raw/dataset.py +0 -223
  37. {litdata-0.2.66 → litdata-0.2.68}/LICENSE +0 -0
  38. {litdata-0.2.66 → litdata-0.2.68}/MANIFEST.in +0 -0
  39. {litdata-0.2.66 → litdata-0.2.68}/requirements.txt +0 -0
  40. {litdata-0.2.66 → litdata-0.2.68}/setup.cfg +0 -0
  41. {litdata-0.2.66 → litdata-0.2.68}/setup.py +0 -0
  42. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/__main__.py +0 -0
  43. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/__init__.py +0 -0
  44. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/commands.py +0 -0
  45. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/__init__.py +0 -0
  46. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/cache.py +0 -0
  47. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/handler/optimize.py +0 -0
  48. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/cli/parser.py +0 -0
  49. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/helpers.py +0 -0
  50. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/imports.py +0 -0
  51. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/__init__.py +0 -0
  52. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/readers.py +0 -0
  53. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/processing/utilities.py +0 -0
  54. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/__init__.py +0 -0
  55. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/raw/types.py +0 -0
  56. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/requirements.py +0 -0
  57. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/__init__.py +0 -0
  58. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/fs_provider.py +0 -0
  59. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/parallel.py +0 -0
  60. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/sampler.py +0 -0
  61. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/serializers.py +0 -0
  62. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/streaming/timing.py +0 -0
  63. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/__init__.py +0 -0
  64. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/_pytree.py +0 -0
  65. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/base.py +0 -0
  66. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/breakpoint.py +0 -0
  67. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/broadcast.py +0 -0
  68. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/dataset_utilities.py +0 -0
  69. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/encryption.py +0 -0
  70. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/env.py +0 -0
  71. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/hf_dataset.py +0 -0
  72. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/packing.py +0 -0
  73. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/parquet.py +0 -0
  74. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/subsample.py +0 -0
  75. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/torch_utils.py +0 -0
  76. {litdata-0.2.66 → litdata-0.2.68}/src/litdata/utilities/train_test_split.py +0 -0
  77. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/dependency_links.txt +0 -0
  78. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/entry_points.txt +0 -0
  79. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/not-zip-safe +0 -0
  80. {litdata-0.2.66 → litdata-0.2.68}/src/litdata.egg-info/top_level.txt +0 -0
@@ -2,7 +2,7 @@
2
2
 
3
3
  Welcome to the PyTorch Lightning community! We're building the most advanced research platform on the planet to implement the latest, best practices and integrations that the amazing PyTorch team and other research organization rolls out!
4
4
 
5
- If you are new to open source, check out [this blog to get started with your first Open Source contribution](https://devblog.pytorchlightning.ai/quick-contribution-guide-86d977171b3a).
5
+ If you are new to open source, check out [GitHub's guide to making your first contribution](https://docs.github.com/en/get-started/quickstart/contributing-to-projects).
6
6
 
7
7
  ## Main Core Value: One less thing to remember
8
8
 
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: litdata
3
- Version: 0.2.66
3
+ Version: 0.2.68
4
4
  Summary: The Deep Learning framework to train, deploy, and ship AI products Lightning fast.
5
5
  Home-page: https://github.com/Lightning-AI/litdata
6
6
  Download-URL: https://github.com/Lightning-AI/litdata
@@ -44,6 +44,7 @@ Requires-Dist: pillow; extra == "extras"
44
44
  Requires-Dist: polars; extra == "extras"
45
45
  Requires-Dist: pyarrow<25.0.0; extra == "extras"
46
46
  Requires-Dist: tqdm; extra == "extras"
47
+ Requires-Dist: uvloop; sys_platform != "win32" and extra == "extras"
47
48
  Requires-Dist: viztracer; extra == "extras"
48
49
  Dynamic: author
49
50
  Dynamic: author-email
@@ -71,12 +72,13 @@ Dynamic: summary
71
72
  &nbsp;
72
73
 
73
74
  <pre>
74
- Transform Optimize
75
+ Transform Optimize / Stream
75
76
 
76
- ✅ Parallelize data processing ✅ Stream large cloud datasets
77
- ✅ Create vector embeddings ✅ Accelerate training by 20x
78
- ✅ Run distributed inference ✅ Pause and resume data streaming
79
- ✅ Scrape websites at scale ✅ Use remote data without local loading
77
+ ✅ Parallelize data processing ✅ Stream raw files with no prep
78
+ ✅ Create vector embeddings ✅ Stream large cloud datasets
79
+ ✅ Run distributed inference ✅ Accelerate training by 20x
80
+ ✅ Scrape websites at scale ✅ Pause and resume data streaming
81
+ ✅ Use remote data without local loading
80
82
  </pre>
81
83
 
82
84
  ---
@@ -92,6 +94,7 @@ Transform Optimize
92
94
  <a href="#speed-up-model-training">Optimize data</a> •
93
95
  <a href="#transform-datasets">Transform data</a> •
94
96
  <a href="#key-features">Features</a> •
97
+ <a href="#stream-raw">Stream raw files</a> •
95
98
  <a href="#resolve-paths">Paths & cloud URLs</a> •
96
99
  <a href="#benchmarks">Benchmarks</a> •
97
100
  <a href="#start-from-a-template">Templates</a> •
@@ -146,6 +149,8 @@ Install all the extras
146
149
  pip install 'litdata[extras]'
147
150
  ```
148
151
 
152
+ On Linux/macOS, `[extras]` includes optional `uvloop` for a faster asyncio event loop used by `StreamingRawDataset` (stdlib asyncio is the fallback when it is not installed).
153
+
149
154
  </details>
150
155
 
151
156
  <details>
@@ -168,29 +173,38 @@ Source: [`.claude/skills/litdata/`](.claude/skills/litdata/) in this repository
168
173
  # Speed up model training
169
174
  Stream datasets directly from cloud storage without local downloads. Choose the approach that fits your workflow:
170
175
 
171
- ## Option 1: Start immediately with existing data ⚡⚡
172
- Stream raw files directly from cloud storage - no pre-optimization needed.
176
+ ## Option 1: Stream existing files as-is ⚡⚡ — `StreamingRawDataset`
177
+
178
+ **No optimize step.** Point LitData at a folder of images, audio, text, or any files (local or cloud) and train with a normal PyTorch `DataLoader`. Downloads are **fully asynchronous** and **batched**; cloud clients include **built-in retries**. You receive **raw `bytes`** — decode, parse, or transform however you want.
179
+
180
+ Details → [Stream raw files](#stream-raw).
173
181
 
174
182
  ```python
175
183
  from litdata import StreamingRawDataset
176
184
  from torch.utils.data import DataLoader
185
+ from PIL import Image
186
+ import io
177
187
 
178
- # Point to your existing cloud data
179
- dataset = StreamingRawDataset("s3://my-bucket/raw-data/")
180
- dataloader = DataLoader(dataset, batch_size=32)
188
+ dataset = StreamingRawDataset(
189
+ "s3://my-bucket/raw-images/", # or gs://, azure://, /teamspace/s3_connections/..., local path
190
+ transform=lambda b: Image.open(io.BytesIO(b)).convert("RGB"), # optional — default is raw bytes
191
+ )
192
+ loader = DataLoader(dataset, batch_size=32, num_workers=8)
181
193
 
182
- for batch in dataloader:
183
- # Process raw bytes on-the-fly
184
- pass
194
+ for batch in loader:
195
+ train_step(batch)
185
196
  ```
186
197
 
187
198
  **Key benefits:**
188
199
 
189
- ✅ **Instant access:** Start streaming immediately without preprocessing.
190
- ✅ **Zero setup time:** No data conversion or optimization required.
191
- ✅ **Native format:** Work with original file formats (images, text, etc.).
192
- ✅ **Flexible processing:** Apply transformations on-the-fly during streaming.
193
- ✅ **Cloud-native:** Stream directly from S3, GCS, or Azure storage.
200
+ ✅ **Zero preprocess:** No chunking job — use the files you already have.
201
+ ✅ **Raw bytes, your rules:** Each sample is file `bytes`; decode with PIL, torchaudio, json, or any custom logic (`transform=` optional).
202
+ ✅ **Fully async + batched:** Concurrent downloads via `asyncio` / `__getitems__` (not one-file-at-a-time).
203
+ ✅ **Built-in retries:** Cloud downloads retry transient failures (adaptive client retries).
204
+ ✅ **Cloud-native:** S3 / GCS / Azure / Studio connections; same path resolver as optimized streaming.
205
+ ✅ **Grouped samples:** Override `setup()` to yield image+mask, audio+transcript, etc.
206
+ ✅ **Indexed once:** `index.json.zstd` cached locally and on the bucket for fast restarts.
207
+ ✅ **Upgrade path:** When I/O becomes the bottleneck, `optimize` → `StreamingDataset` for max throughput.
194
208
 
195
209
  ## Option 2: Optimize for maximum performance ⚡⚡⚡
196
210
  Accelerate model training (20x faster) by optimizing datasets for streaming directly from cloud storage. Work with remote data without local downloads with features like loading data subsets, accessing individual samples, and resumable streaming.
@@ -227,7 +241,7 @@ if __name__ == "__main__":
227
241
  inputs=list(range(1000)), # the inputs to the function (here it's a list of numbers)
228
242
  output_dir="fast_data", # optimized data is stored here
229
243
  num_workers=4, # the number of workers on the same machine
230
- chunk_bytes="64MB" # size of each chunk
244
+ chunk_bytes="64MB" # default; see FAQ for larger samples
231
245
  )
232
246
  ```
233
247
 
@@ -265,6 +279,22 @@ for sample in dataloader:
265
279
  img, cls = sample["image"], sample["class"]
266
280
  ```
267
281
 
282
+ **Keyed lookup and in-place patches** (needs `polars` and `optimize(..., key_fn=...)` or `build_keys_index`):
283
+
284
+ ```python
285
+ ld.optimize(fn=fn, inputs=inputs, output_dir="fast_data", chunk_bytes="64MB", key_fn=lambda s: s["id"])
286
+
287
+ ds = ld.StreamingDataset("fast_data")
288
+ sample = ds["entity-id"] # str keys
289
+ sample = ds.get_by_key(42) # int entity keys; ds[42] is still positional
290
+
291
+ with ld.dataset_update("fast_data") as update: # local directory only
292
+ update["entity-id"] = {"id": "entity-id", "x": 1}
293
+ update.commit()
294
+ ```
295
+
296
+ `mode="append"` continues chunk numbering. `use_checkpoint=True` tries to resume an interrupted optimize. They are not the same.
297
+
268
298
  **Key benefits:**
269
299
 
270
300
  ✅ **Accelerate training:** Optimized datasets load 20x faster.
@@ -325,79 +355,158 @@ ld.map(
325
355
  ## Features for optimizing and streaming datasets for model training
326
356
 
327
357
  <details>
328
- <summary> ✅ Stream raw datasets from cloud storage (beta) <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
358
+ <summary> ✅ Stream raw files as-is (no optimize) — StreamingRawDataset <a id="stream-raw" href="#stream-raw">🔗</a> </summary>
329
359
  &nbsp;
330
360
 
331
- Effortlessly stream raw files (images, text, etc.) directly from S3, GCS, and Azure cloud storage without any optimization or conversion. Ideal for workflows requiring instant access to original data in its native format.
361
+ `StreamingRawDataset` streams **your existing files** from local disk or cloud storage with **no conversion step**. It is a map-style `torch.utils.data.Dataset`: use a standard PyTorch `DataLoader` (not `StreamingDataLoader`).
332
362
 
333
- **Prerequisites:**
363
+ **You get raw `bytes`.** LitData does not impose a sample schema — open images with PIL, parse JSONL, decode audio, run your own tokenizer, or pass a `transform=` if you prefer. Grouped items yield `list[bytes]` (e.g. image + mask).
334
364
 
335
- Install the required dependencies to stream raw datasets from cloud storage like **Amazon S3** or **Google Cloud Storage**:
365
+ Downloads are **fully asynchronous** and **batched**: when the DataLoader requests a batch, `__getitems__` fetches those files concurrently with `asyncio.gather`. Cloud clients include **built-in retries** for transient network errors.
336
366
 
337
- ```bash
338
- # for aws s3
339
- pip install "litdata[extra]" s3fs
367
+ Use it when you want to train or prototype on JPEGs, masks, audio, JSONL, etc. **immediately**. Switch to [`optimize` → `StreamingDataset`](#speed-up-model-training) later if you need maximum cloud training throughput.
368
+
369
+ | | `StreamingRawDataset` | `StreamingDataset` (optimized) |
370
+ |--|----------------------|--------------------------------|
371
+ | Prep | None — point at a folder | One-time `optimize` → `chunk-*.bin` + `index.json` |
372
+ | Item | **Raw file `bytes`** (you decide how to decode) | Deserialized samples (dict/tensor/…) |
373
+ | I/O | Fully async, batched downloads + retries | Chunk prefetch / cache pipeline |
374
+ | Loader | `torch.utils.data.DataLoader` | Prefer `StreamingDataLoader` (shuffle, resume) |
375
+ | Best for | Instant start, full control over bytes | Highest sustained training I/O |
340
376
 
341
- # for gcloud storage
342
- pip install "litdata[extra]" gcsfs
377
+ ### Install (cloud)
378
+
379
+ ```bash
380
+ pip install "litdata[extra]" s3fs # Amazon S3
381
+ pip install "litdata[extra]" gcsfs # Google Cloud Storage
382
+ # Azure / Studio connections: see Paths & cloud URLs
343
383
  ```
344
384
 
345
- **Usage Example:**
385
+ ### Quick start
386
+
346
387
  ```python
347
388
  from torch.utils.data import DataLoader
348
389
  from litdata import StreamingRawDataset
390
+ from PIL import Image
391
+ import io
349
392
 
350
- dataset = StreamingRawDataset("s3://bucket/files/")
393
+ def to_image(data: bytes):
394
+ return Image.open(io.BytesIO(data)).convert("RGB")
395
+
396
+ dataset = StreamingRawDataset(
397
+ "s3://my-bucket/images/", # also: gs://, azure://, /teamspace/s3_connections/..., local path
398
+ transform=to_image, # optional; default yields raw bytes
399
+ storage_options={}, # optional cloud credentials / endpoint
400
+ )
401
+ loader = DataLoader(dataset, batch_size=32, num_workers=8)
351
402
 
352
- # Use with PyTorch DataLoader
353
- loader = DataLoader(dataset, batch_size=32)
354
403
  for batch in loader:
355
- # Each item is raw bytes
356
- pass
404
+ train_step(batch)
357
405
  ```
358
406
 
359
- > Use `StreamingRawDataset` to stream your data as-is. Use `StreamingDataset` for fastest streaming after optimizing your data.
407
+ ### Constructor knobs
408
+
409
+ | Arg | Default | Purpose |
410
+ |-----|---------|---------|
411
+ | `input_dir` | required | Folder URL/path (same [resolver](#resolve-paths) as optimized streaming) |
412
+ | `cache_dir` | LitData default cache | Where the file index (and optional file cache) live |
413
+ | `cache_files` | `False` | If `True`, keep downloaded files on disk under `cache_dir` (mirror remote layout) |
414
+ | `recompute_index` | `False` | Force re-scan when remote files changed |
415
+ | `transform` | `None` | `fn(bytes) -> Any` or `fn(list[bytes]) -> Any` for grouped items |
416
+ | `storage_options` | `{}` | Cloud client options |
417
+ | `indexer` | `FileIndexer()` | Custom discovery (subclass `BaseIndexer`) |
418
+ | `max_concurrent_downloads` | `None` (adaptive) | Per-worker in-flight downloads. `None` = size-aware budget (bandwidth; Little’s-law only for medians &lt;~8 MiB) split across workers; single-process capped at 128. An explicit `int` is used exactly (no silent clamp) |
419
+ | `max_prefetch` | `16` | Per-worker sequential look-ahead after each batch (default on). When `num_workers > 1`, effective look-ahead is `min(max_prefetch, 64 // num_workers)` so aggregate stays ~64 items. Pass `0` to disable |
420
+ | `prefetch_cache_size` | auto | LRU cap for prefetched items (defaults from `max_prefetch`) |
421
+ | `hedge_delay` | `0` | Seconds before a hedged duplicate GET for a slow download (`0` = off, default; opt-in) |
422
+ | `range_parallel_threshold` | `0` | Objects ≥ this many bytes use parallel ranged GETs (`0` = whole-object only; opt-in) |
423
+ | `item_type` | `"bytes"` | `"bytes"` buffers in RAM; `"path"` returns local cache paths (`cache_files=True` required) |
360
424
 
425
+ ### Group related files (`setup`)
361
426
 
362
- You can also customize how files are grouped by subclassing `StreamingRawDataset` and overriding the `setup` method. This is useful for pairing related files (e.g., image and mask, audio and transcript) or any custom grouping logic.
427
+ Default: **one file = one sample**. Override `setup` to filter or group (image + mask, audio + transcript, …). Return either a list of `FileMetadata` or a list of groups (`list[list[FileMetadata]]`).
363
428
 
364
429
  ```python
365
- from typing import Union
430
+ from collections import defaultdict
366
431
  from torch.utils.data import DataLoader
367
432
  from litdata import StreamingRawDataset
368
433
  from litdata.raw.indexer import FileMetadata
369
434
 
370
435
  class SegmentationRawDataset(StreamingRawDataset):
371
- def setup(self, files: list[FileMetadata]) -> Union[list[FileMetadata], list[list[FileMetadata]]]:
372
- # TODO: Implement your custom grouping logic here.
373
- # For example, group files by prefix, extension, or any rule you need.
374
- # Return a list of groups, where each group is a list of FileMetadata.
375
- # Example:
376
- # return [[image, mask], ...]
377
- pass
378
-
379
- # Initialize the custom dataset
380
- dataset = SegmentationRawDataset("s3://bucket/files/")
381
- loader = DataLoader(dataset, batch_size=32)
382
- for item in loader:
383
- # Each item in the batch is a pair: [image_bytes, mask_bytes]
384
- pass
436
+ def setup(self, files: list[FileMetadata]) -> list[list[FileMetadata]]:
437
+ # Pair img_001.jpg with img_001.png (mask) by stem
438
+ by_stem: dict[str, dict[str, FileMetadata]] = defaultdict(dict)
439
+ for f in files:
440
+ name = f.path.rsplit("/", 1)[-1]
441
+ stem, _, ext = name.rpartition(".")
442
+ by_stem[stem][ext.lower()] = f
443
+ items = []
444
+ for stem, parts in sorted(by_stem.items()):
445
+ if "jpg" in parts and "png" in parts:
446
+ items.append([parts["jpg"], parts["png"]])
447
+ return items
448
+
449
+ dataset = SegmentationRawDataset(
450
+ "s3://bucket/seg/",
451
+ transform=lambda pair: (pair[0], pair[1]), # list[bytes]: [image, mask]
452
+ )
453
+ loader = DataLoader(dataset, batch_size=16, num_workers=4)
454
+ for images, masks in loader:
455
+ ...
385
456
  ```
386
457
 
387
- **Smart Index Caching**
458
+ ### Index caching (`index.json.zstd`)
388
459
 
389
- `StreamingRawDataset` automatically caches the file index for instant startup. Initial scan, builds and caches the index, then subsequent runs load instantly.
460
+ First open scans the tree and writes a compressed file list:
390
461
 
391
- **Two-Level Cache:**
392
- - **Local:** Stored in your cache directory for instant access
393
- - **Remote:** Automatically saved to cloud storage (e.g., `s3://bucket/files/index.json.zstd`) for reuse
462
+ - **Local cache** under your LitData cache dir (fast restart on the same machine)
463
+ - **Remote copy** next to the data when possible (e.g. `s3://bucket/files/index.json.zstd`) so every machine skips the scan
394
464
 
395
- **Force Rebuild:**
396
465
  ```python
397
- # When dataset files have changed
466
+ # After adding/removing files on the bucket:
398
467
  dataset = StreamingRawDataset("s3://bucket/files/", recompute_index=True)
399
468
  ```
400
469
 
470
+ Do **not** confuse this with optimized LitData’s `index.json` (chunk metadata). Raw indexing only lists files.
471
+
472
+ ### How downloads work
473
+
474
+ 1. DataLoader asks for a batch of indices → `__getitems__`.
475
+ 2. LitData **asynchronously** downloads those files **in parallel** (`asyncio.gather` + `adownload_fileobj`).
476
+ 3. Cloud SDKs apply **retries** on transient failures (e.g. S3 adaptive retries).
477
+ 4. Each item is returned as **`bytes`** (or `list[bytes]` if `setup` grouped files), then optional `transform`.
478
+
479
+ Your training loop stays normal PyTorch — no async/`await` in user code.
480
+
481
+ ```python
482
+ # Default: you own the bytes
483
+ dataset = StreamingRawDataset("s3://bucket/files/")
484
+ raw: bytes = dataset[0]
485
+ # e.g. Image.open(io.BytesIO(raw)), json.loads(raw), np.frombuffer(raw), ...
486
+ ```
487
+
488
+ ### Tips
489
+
490
+ - Prefer `num_workers > 0` so worker processes overlap async batch downloads with training. Scale workers toward host vCPUs for network-bound JPEG-sized objects — avoid saturating every vCPU.
491
+ - On Linux, after any parent-process dataset I/O, use `DataLoader(..., multiprocessing_context="spawn", persistent_workers=True)` — default `fork` can hang S3 clients in workers.
492
+ - Default `max_prefetch=16` enables sequential look-ahead **per DataLoader worker**; shuffled access disables it. Pass `0` to turn off. When `num_workers > 1`, look-ahead and download concurrency both scale down with worker count so aggregate in-flight work stays bounded.
493
+ - Prefer an `s3://` / `gs://` URL or `/teamspace/s3_connections/...` so LitData hits the bucket directly ([resolver](#resolve-paths)) — avoid reading through FUSE.
494
+ - Leave `range_parallel_threshold=0` (default) for typical JPEGs; raise it only for large objects where parallel ranged GETs help.
495
+ - Best for medium/large files. Tiny objects (≲100 KB) are request-overhead bound — pack with [`optimize`](#speed-up-model-training) → `StreamingDataset` when I/O plateaus.
496
+
497
+ ### Throughput
498
+
499
+ On ImageNet val raw over S3 (50 k JPEGs, batch size 64, spawn workers), throughput gains are clearest at **low worker counts / notebooks** (**+20–80%** at ≤8 workers). At **high workers** (≥16), results are roughly **parity within run-to-run noise**.
500
+
501
+ | workers | before | after | Δ |
502
+ |--------:|-------:|------:|--:|
503
+ | 0 | 543 | 735 | **+35%** |
504
+ | 2 | 816 | 1475 | **+81%** |
505
+ | 8 | 4841 | 5718 | **+18%** |
506
+ | 16+ | ~6k | ~6k | ~parity |
507
+
508
+ Useful knobs: `num_workers`, `max_prefetch` (default 16; worker-aware), `download_timeout` (batch-level hang protection). Ranged parallel downloads stay opt-in (`range_parallel_threshold=0`).
509
+
401
510
  </details>
402
511
 
403
512
  <details>
@@ -692,6 +801,10 @@ Shuffling is **deterministic** and designed for distributed training:
692
801
 
693
802
  The permutation depends on `seed`, the epoch, and chunk metadata — the same settings always yield the same order (required for resumable `state_dict`).
694
803
 
804
+ **Object storage (`s3://`, `gs://`, …)** globally permutes chunks (`FullShuffle`). Random chunk order is cheap once files are already copied into the local cache.
805
+
806
+ **POSIX-fast** (automatic for any local path) mmaps chunks in place. **Vast / NFS / Lustre / GPFS** (and `LITDATA_POSIX_FAST=1`) use `WindowShuffle`: each worker gets **whole chunks** in a sequential stripe, then shuffles only inside a sliding window (default **16**, `LITDATA_POSIX_SHUFFLE_WINDOW`) for both chunk order and in-chunk items. Local disks (ext4/xfs) keep global `FullShuffle`. Object URLs stay on `FullShuffle`. `LITDATA_POSIX_FAST=0` disables in-place mmap.
807
+
695
808
  ```python
696
809
  from litdata import StreamingDataset, StreamingDataLoader
697
810
 
@@ -715,6 +828,35 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
715
828
 
716
829
  </details>
717
830
 
831
+ <details>
832
+ <summary> ✅ FAQ: chunk size &amp; shuffle before optimize <a id="faq-chunk-shuffle" href="#faq-chunk-shuffle">🔗</a> </summary>
833
+ &nbsp;
834
+
835
+ ### What `chunk_bytes` should I use?
836
+
837
+ Default is **64MB** — a good starting point for typical small/medium samples.
838
+
839
+ When each datapoint is large (e.g. a few MB), prefer a **larger chunk** (practical range often **256–512MB**) so each chunk holds more samples and **intra-chunk batch randomization** has a bigger pool. Tradeoff: larger chunks take **longer to download** before they can be used.
840
+
841
+ This is expert guidance (recommended-range mindset), not a published chunk-size sweep.
842
+
843
+ ### Is StreamingDataset shuffle enough if my source data is ordered?
844
+
845
+ **Not always.** LitData handles **distributed sampling** and **bucket sampling within chunks** automatically (`shuffle=True` randomizes chunk order and item order inside each chunk). That is **not** a substitute for a fully shuffled file-level DataLoader when the source has strong structure (same subject/set contiguous, class blocks, etc.).
846
+
847
+ If ordered data would make chunked sampling problematic and you cannot embed the grouping as the sample unit:
848
+
849
+ - Shuffle the list of samples **before** `optimize` so chunks mix well, **or**
850
+ - Use [`StreamingRawDataset`](#stream-raw) (per-file random access via a standard PyTorch `DataLoader` with `shuffle=True`) instead of optimize → `StreamingDataset`.
851
+
852
+ ### FUSE vs LitData (Lightning Studios)
853
+
854
+ `/teamspace/s3_connections` (and related mounts) are **FUSE** — fine for browsing, not for training I/O. Under load they are very slow and can crash. Pass the same path into LitData (`StreamingRawDataset` / `StreamingDataset` / `optimize`): LitData resolves it and talks **directly** to the bucket ([Resolve any path](#resolve-paths)).
855
+
856
+ Rough ImageNet order-of-magnitude on a Studio (not hard guarantees; right tuning for raw): FUSE hand-read ~**600** images/s · [`StreamingRawDataset`](#stream-raw) ~**6–7k** · optimized [`StreamingDataset`](#speed-up-model-training) (64MB chunks) ~**11k**.
857
+
858
+ </details>
859
+
718
860
  <details>
719
861
  <summary> ✅ StreamingDataset & StreamingDataLoader knobs <a id="streaming-kwargs" href="#streaming-kwargs">🔗</a> </summary>
720
862
  &nbsp;
@@ -731,7 +873,7 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
731
873
  | `seed` | `42` | Shuffle / subsample RNG |
732
874
  | `serializers` | built-ins | Custom serialize/deserialize map |
733
875
  | `max_cache_size` | `"100GB"` | Evict consumed chunks beyond this size |
734
- | `max_pre_download` | `2` | Chunks each worker may prefetch (raise for throughput; watch disk) |
876
+ | `max_pre_download` | `2` | Chunks each worker may prefetch (raise for throughput; watch disk / RAM) |
735
877
  | `subsample` | `1.0` | Fraction of data (`0.01`) or upsample (`2.5`) |
736
878
  | `encryption` | `None` | `FernetEncryption` / `RSAEncryption` / custom |
737
879
  | `storage_options` | `{}` | Cloud client options |
@@ -742,6 +884,8 @@ loader = StreamingDataLoader(train, batch_size=64, shuffle=True, drop_last=True)
742
884
 
743
885
  Peak disk ≈ `num_workers × max_pre_download × mean_chunk_size`.
744
886
 
887
+ On **Vast / NFS / local disk**, POSIX-fast is on by default (`LITDATA_POSIX_FAST=0` to disable). `WILLNEED` prefetch and `num_workers` are capped when they would exceed about half of `MemAvailable`. Idle **hugepages** (common on GPU nodes) do not count as available RAM — drop unused `nr_hugepages` if `MemAvailable` looks tiny next to `MemTotal`.
888
+
745
889
  **`StreamingDataLoader`**
746
890
 
747
891
  | Argument | Description |
@@ -1093,7 +1237,7 @@ if __name__ == "__main__":
1093
1237
 
1094
1238
  Mix and match different sets of data to experiment and create better models.
1095
1239
 
1096
- Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://huggingface.co/datasets/cerebras/SlimPajama-627B) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
1240
+ Combine datasets with `CombinedStreamingDataset`. As an example, this mixture of [Slimpajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) was used in the [TinyLLAMA](https://github.com/jzhang38/TinyLlama) project to pretrain a 1.1B Llama model on 3 trillion tokens.
1097
1241
 
1098
1242
  ```python
1099
1243
  from litdata import StreamingDataset, CombinedStreamingDataset, StreamingDataLoader, TokensLoader
@@ -1684,7 +1828,7 @@ Only **worker 0** is instrumented. When an `int` is used, the tracer wraps `fetc
1684
1828
 
1685
1829
  - Delete or change `profile_dir` between runs — LitData removes an existing `result.json` before starting.
1686
1830
  - Pair with a wiped chunk cache if you care about **cold** epoch behavior (`litdata cache clear`).
1687
- - For deeper LitData internals (download / lock / delete timeline), use `enable_tracer()` + [Litracer](https://github.com/deependujha/litracer) instead — see [Debug & Profile LitData](#debug-profile). That path is complementary: viztracer = DataLoader worker CPU timeline; Litracer = LitData pipeline events.
1831
+ - For deeper LitData internals (download / read / delete timeline), use `enable_tracer()` + [Litracer](https://github.com/Lightning-AI/litracer) instead — see [Debug & Profile LitData](#debug-profile). That path is complementary: viztracer = DataLoader worker CPU timeline; Litracer = LitData pipeline events.
1688
1832
 
1689
1833
  </details>
1690
1834
 
@@ -1797,6 +1941,9 @@ export LITDATA_ASYNC_MIN_PRE_DOWNLOAD=0
1797
1941
  | `LITDATA_DISABLE_VERSION_CHECK` | `0` | `1` skips the upgrade tip |
1798
1942
  | `HF_TOKEN` | — | Gated Hugging Face datasets |
1799
1943
  | `DEBUG_LITDATA` / `PRINT_DEBUG_LOGS` | `0` | Internal debug / stdout logs |
1944
+ | `LITDATA_LOG_FILE` | `litdata_debug.log` | `enable_tracer()` output path |
1945
+ | `LITDATA_TRACE_LEVEL` | unset | `batch` / `chunk` / `sample` / `debug` / `off` (see [Debug & Profile](#debug-profile)) |
1946
+ | `LITDATA_TRACE_CATEGORIES` | from level | Comma-separated cats, e.g. `download,read,delete` |
1800
1947
 
1801
1948
  Multi-node `optimize`/`map` on Studios also uses `DATA_OPTIMIZER_*` (set by the platform). Full catalog (debug logs, Studio injects, torchrun): see the LitData skill `reference/env-vars.md` when using agent skills, or the source modules `constants.py` / `async_prefetch.py`.
1802
1949
 
@@ -1962,66 +2109,68 @@ ds = StreamingDataset(input_dir=data_dir, encryption=rsa)
1962
2109
 
1963
2110
  &nbsp;
1964
2111
 
1965
- LitData comes with built-in logging and profiling capabilities to help you debug and profile your data streaming workloads.
2112
+ `enable_tracer()` records the streaming pipeline (download vs read vs delete vs batch) as one-line events. [Litracer](https://github.com/Lightning-AI/litracer) converts that log into a Chrome / [Perfetto](https://ui.perfetto.dev) trace.
1966
2113
 
1967
- <img width="1439" alt="431247797-0e955e71-2f9a-4aad-b7c1-a8218fed2e2e" src="https://github.com/user-attachments/assets/4e40676c-ba0b-49af-acac-975977173669" />
2114
+ This is complementary to [`profile_batches`](#profile-loading) (viztracer = DataLoader worker **CPU**; Litracer = LitData **pipeline** events).
1968
2115
 
1969
- - e.g., with LitData Streaming
2116
+ <img width="1439" alt="431247797-0e955e71-2f9a-4aad-b7c1-a8218fed2e2e" src="https://github.com/user-attachments/assets/4e40676c-ba0b-49af-acac-975977173669" />
1970
2117
 
1971
2118
  ```python
1972
2119
  import litdata as ld
1973
2120
  from litdata.debugger import enable_tracer
1974
2121
 
1975
- # WARNING: Remove existing trace `litdata_debug.log` file if it exists before re-tracing
1976
- enable_tracer()
2122
+ # Call once per process, before the DataLoader. Delete an existing log before re-tracing (append).
2123
+ enable_tracer(level="chunk", log_file="litdata_debug.log")
2124
+ # level="batch" | "chunk" (default) | "sample" | "debug" | "off"
2125
+ # enable_tracer(categories=["download", "read", "delete"])
1977
2126
 
1978
2127
  if __name__ == "__main__":
1979
2128
  dataset = ld.StreamingDataset("s3://my-bucket/my-data", shuffle=True)
1980
- dataloader = ld.StreamingDataLoader(dataset, batch_size=64)
1981
-
1982
- for batch in dataloader:
1983
- print(batch) # Replace with your data processing logic
2129
+ for batch in ld.StreamingDataLoader(dataset, batch_size=64, num_workers=8):
2130
+ ...
1984
2131
  ```
1985
2132
 
1986
- 1. Generate Debug Log:
2133
+ | Level | Events |
2134
+ | ----- | ------ |
2135
+ | `batch` | Epoch + per-batch spans, plus crashes |
2136
+ | `chunk` (default) | + `download`, `read`, `delete`, `decompress`, `prefetch` |
2137
+ | `sample` | + per-item `__getitem__` (high volume) |
2138
+ | `debug` | + `.cnt` lock refcount spans |
2139
+ | `off` | Disable |
1987
2140
 
1988
- - Run your Python program and it'll create a log file containing detailed debug information.
2141
+ Event **names** are stable (`download`, `read`, `delete`, `batch`, `sample`, `crash`). Chunk / sample indexes live in args so Perfetto groups all downloads together. Each line is `key: value;` pairs with Chrome **microsecond** timestamps. Crashes are a one-line instant (`ph: I`, `name: crash`); the Python traceback is printed to **stderr**, not the log file (a multi-line `logger.exception` would break Litracer).
1989
2142
 
1990
- ```bash
1991
- python main.py
1992
- ```
2143
+ Env overrides: `LITDATA_LOG_FILE`, `LITDATA_TRACE_LEVEL`, `LITDATA_TRACE_CATEGORIES` (comma-separated). Tracer calls are no-ops when tracing is off.
1993
2144
 
1994
- 2. Install [Litracer](https://github.com/deependujha/litracer/):
2145
+ 1. Generate the log:
1995
2146
 
1996
- - Option 1: Using Go (recommended)
1997
- - Install Go on your system.
1998
- - Run the following command to install Litracer:
2147
+ ```bash
2148
+ python train.py # writes litdata_debug.log
2149
+ ```
1999
2150
 
2000
- ```bash
2001
- go install github.com/deependujha/litracer@latest
2002
- ```
2151
+ 2. Install [Litracer](https://github.com/Lightning-AI/litracer) (Go 1.23+):
2003
2152
 
2004
- - Option 2: Download Binary
2005
- - Visit the [LitRacer GitHub Releases](https://github.com/deependujha/litracer/releases) page.
2006
- - Download the appropriate binary for your operating system and follow the installation instructions.
2153
+ ```bash
2154
+ git clone https://github.com/Lightning-AI/litracer.git
2155
+ cd litracer && go build -o litracer .
2156
+ ```
2007
2157
 
2008
- 3. Convert Debug Log to trace JSON:
2158
+ Or `go install github.com/deependujha/litracer@latest` (published Go module path). Until `go.mod` is renamed, `go install github.com/Lightning-AI/litracer@latest` does not work. Release binaries: [GitHub Releases](https://github.com/Lightning-AI/litracer/releases).
2009
2159
 
2010
- - Use litracer to convert the generated log file into a trace JSON file. This command uses 100 workers for conversion:
2160
+ 3. Convert and open in Perfetto:
2011
2161
 
2012
2162
  ```bash
2013
- litracer litdata_debug.log -o litdata_trace.json -w 100
2163
+ litracer --quiet --validate -o litdata_trace.json.gz litdata_debug.log
2164
+ litracer --quiet --cat download,read,delete -o io.json.gz litdata_debug.log
2165
+ # open the .json.gz at https://ui.perfetto.dev (preferred) or chrome://tracing
2014
2166
  ```
2015
2167
 
2016
- 4. Visualize the trace:
2017
-
2018
- - Use either `chrome://tracing` in the Chrome browser or `ui.perfetto.dev` to view the `litdata_trace.json` file for in-depth performance insights. You can also use `SQL queries` to analyze the logs.
2019
- - `Perfetto` is recommended over `chrome://tracing` for visualization & analyzing.
2168
+ `--quiet` prints a one-line summary (per-category durations, unmatched B/E, crashes). `--cat` keeps only those categories. Matched B/E pairs become complete (`ph: X`) spans unless `--no-complete`. Default output is gzip Chrome JSON (`.json.gz`) — both Perfetto and `chrome://tracing` open it; pass `-o file.json` for uncompressed.
2020
2169
 
2021
- - Key Points:
2170
+ - For trace files `> 2GB`, see [Perfetto large traces](https://perfetto.dev/docs/visualization/large-traces).
2171
+ - If you connect Perfetto to the RPC server, prefer Chrome over Brave (Brave often does not autodetect the RPC server).
2022
2172
 
2023
- - For very large trace.json files (`> 2GB`), refer to the [Perfetto documentation](https://perfetto.dev/docs/visualization/large-traces) for using native accelerators.
2024
- - If you are trying to connect Perfetto to the RPC server, it is recommended to use Chrome over Brave, as it has been observed that Perfetto in Brave does not autodetect the RPC server.
2173
+ **Multi-worker `s3://` `FileNotFoundError` after ~120s:** `num_workers=0` working while `num_workers>0` fails usually means the DataLoader parent started obstore (tokio) before fork and worker GETs hung. Current LitData fetches `index.json` with boto3 so workers can lazy-init obstore; they fall back to boto3 if the parent already started the runtime. On Studio R2 / `lightning_storage`, the same symptom can be a prefetch-thread crash (`data_connection_id` / `endpoint_url` into `boto3.Session`) — look for `[litdata] PrepareChunksThread CRASHED` on stderr and a `crash` instant in the trace.
2025
2174
 
2026
2175
  </details>
2027
2176
 
@@ -2213,7 +2362,7 @@ Full knob list for `litdata.optimize` (see Quick start for the minimal recipe).
2213
2362
  | `output_dir` | `"optimized_data"` | Local or cloud ([resolver](#resolve-paths)); version remote prefixes |
2214
2363
  | `input_dir` | `None` | Remote input root for background download |
2215
2364
  | `weights` | `None` | Per-input weights to balance workers |
2216
- | `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`) |
2365
+ | `chunk_bytes` | `None` | Max bytes per chunk (e.g. `"64MB"`; see [FAQ](#faq-chunk-shuffle) for larger samples) |
2217
2366
  | `chunk_size` | `None` | Max items (or tokens with `TokensLoader`) per chunk |
2218
2367
  | `align_chunking` | `False` | Match single-worker chunk boundaries (needs `chunk_size`; uneven load) |
2219
2368
  | `compression` | `None` | `"zstd"` today |
@@ -2306,6 +2455,17 @@ Speed to stream Imagenet 1.2M from local disk with ffcv vs LitData:
2306
2455
  | ffcv(os_cache=True) | JPEG 90% | 20 GB | 7653 | 8051 |
2307
2456
  | ffcv(os_cache=False) | JPEG 90% | 20 GB | 8149 | 8607 |
2308
2457
 
2458
+ Speed to stream a **synthetic ImageNet-scale set from Vast NFS** (NFSv3 `nconnect=32`, 208-CPU host, ~1 TiB RAM). Dataset: **1.08M** JPEG q95 256×256 (~160 GiB, 64 MiB chunks). `StreamingDataLoader`, batch **256**, `shuffle=True`, `drop_last=True`, decode only unless noted. POSIX-fast mmaps chunks **in place** (no copy into `~/.lightning/chunks`).
2459
+
2460
+ | Setup | Workers | Images / sec |
2461
+ |---|---|---|
2462
+ | Copy into local cache (`LITDATA_POSIX_FAST=0`) | 48 | **16.7k** (2-epoch avg) |
2463
+ | POSIX-fast (this default on local/Vast paths) | 48 | **18.2k** |
2464
+ | POSIX-fast + README ImageNet augs (crop 224, flip, float32) | 48 | **12.9k** |
2465
+ | POSIX-fast, all CPU cores | **208** | **35.8k** |
2466
+
2467
+ Notes: 208 workers need enough **MemAvailable**. This host had **928×1 GiB hugepages** reserved and idle (~900 GiB locked); after `nr_hugepages=0`, 208 workers stayed healthy. If `num_workers=os.cpu_count()` would crowd RAM, LitData **clamps** workers (`LITDATA_POSIX_MAX_WORKERS=0` disables) and skips `WILLNEED` prefetch. Real ImageNet JPEG 90% is much smaller (~12 GiB) and usually decodes faster than this q95 noise set.
2468
+
2309
2469
  ### Raw Dataset
2310
2470
 
2311
2471
  Speed to stream raw Imagenet 1.2M from different cloud storage providers:
@@ -2316,8 +2476,7 @@ Speed to stream raw Imagenet 1.2M from different cloud storage providers:
2316
2476
  | AWS S3 | ~6400 +/- 100 | ~3200 +/- 100 |
2317
2477
  | Google Cloud Storage | ~5650 +/- 100 | ~3100 +/- 100 |
2318
2478
 
2319
- > **Note:**
2320
- > Use `StreamingRawDataset` if you want to stream your data as-is. Use `StreamingDataset` if you want the fastest streaming and are okay with optimizing your data first.
2479
+ > **Also see:** [`StreamingRawDataset`](#stream-raw) streams existing files with **no optimize step** (great default to start). Use `StreamingDataset` after `optimize` when you need the highest sustained training throughput.
2321
2480
 
2322
2481
  &nbsp;
2323
2482
 
@@ -2406,7 +2565,7 @@ Below are templates for real-world applications of LitData at scale.
2406
2565
  | -------------------------------- | ----------------- | ----------------- | -------------- | -------------- |
2407
2566
  | [Benchmark cloud data-loading libraries](https://lightning.ai/lightning-ai/studios/benchmark-cloud-data-loading-libraries) | Image & Label | 10 | 1 | [Imagenet 1M](https://paperswithcode.com/sota/image-classification-on-imagenet?tag_filter=171) |
2408
2567
  | [Optimize GeoSpatial data for model training](https://lightning.ai/lightning-ai/studios/convert-spatial-data-to-lightning-streaming) | Image & Mask | 120 | 32 | [Chesapeake Roads Spatial Context](https://github.com/isaaccorley/chesapeakersc) |
2409
- | [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://huggingface.co/datasets/cerebras/SlimPajama-627B) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
2568
+ | [Optimize TinyLlama 1T dataset for training](https://lightning.ai/lightning-ai/studios/prepare-the-tinyllama-1t-token-dataset) | Text | 240 | 32 | [SlimPajama](https://www.cerebras.ai/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama) & [StarCoder](https://huggingface.co/datasets/bigcode/starcoderdata) |
2410
2569
  | [Optimize parquet files for model training](https://lightning.ai/lightning-ai/studios/convert-parquets-to-lightning-streaming) | Parquet Files | 12 | 16 | Randomly Generated data |
2411
2570
 
2412
2571
  &nbsp;