@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,482 +0,0 @@
1
- ---
2
- name: dask
3
- description: Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.
4
- allowed-tools: Read Write Edit Bash
5
- license: BSD-3-Clause license
6
- compatibility: Requires Python 3.10+ and dask 2025.1+. DataFrame workflows need pandas 2+ and PyArrow 16+. Cloud paths (s3://, gcs://) need s3fs or gcsfs. Cluster deployment uses dask.distributed (included with dask[complete]).
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # Dask
13
-
14
- ## Overview
15
-
16
- Dask is a Python library for parallel and distributed computing that enables three critical capabilities:
17
- - **Larger-than-memory execution** on single machines for data exceeding available RAM
18
- - **Parallel processing** for improved computational speed across multiple cores
19
- - **Distributed computation** supporting terabyte-scale datasets across multiple machines
20
-
21
- Dask scales from laptops (processing ~100 GiB) to clusters (processing ~100 TiB) while maintaining familiar Python APIs.
22
-
23
- **Current upstream:** dask **2026.3.0** (PyPI, March 2026). Docs: [docs.dask.org](https://docs.dask.org/en/stable/). Since **2025.1.0**, the expression-based DataFrame API with query planning is the only implementation — do not install `dask-expr` separately or set `dataframe.query-planning: False`.
24
-
25
- ## Quick Start
26
-
27
- ### Installation
28
-
29
- ```bash
30
- uv pip install "dask>=2025.1"
31
- ```
32
-
33
- For a typical pandas/NumPy workflow with the distributed scheduler and dashboard:
34
-
35
- ```bash
36
- uv pip install "dask[complete]"
37
- ```
38
-
39
- Remote object storage (S3, GCS, Azure):
40
-
41
- ```bash
42
- uv pip install s3fs # s3:// paths
43
- uv pip install gcsfs # gs:// paths
44
- ```
45
-
46
- Requires **Python 3.10+** (3.9 support dropped in 2024.12). DataFrame I/O requires **PyArrow 16+** (as of dask 2026.1.2).
47
-
48
- ## When to Use This Skill
49
-
50
- This skill should be used when:
51
- - Process datasets that exceed available RAM
52
- - Scale pandas or NumPy operations to larger datasets
53
- - Parallelize computations for performance improvements
54
- - Process multiple files efficiently (CSVs, Parquet, JSON, text logs)
55
- - Build custom parallel workflows with task dependencies
56
- - Distribute workloads across multiple cores or machines
57
-
58
- ## Core Capabilities
59
-
60
- Dask provides five main components, each suited to different use cases:
61
-
62
- ### 1. DataFrames - Parallel Pandas Operations
63
-
64
- **Purpose**: Scale pandas operations to larger datasets through parallel processing.
65
-
66
- **When to Use**:
67
- - Tabular data exceeds available RAM
68
- - Need to process multiple CSV/Parquet files together
69
- - Pandas operations are slow and need parallelization
70
- - Scaling from pandas prototype to production
71
-
72
- **Reference Documentation**: For comprehensive guidance on Dask DataFrames, refer to `references/dataframes.md` which includes:
73
- - Reading data (single files, multiple files, glob patterns)
74
- - Common operations (filtering, groupby, joins, aggregations)
75
- - Custom operations with `map_partitions`
76
- - Performance optimization tips
77
- - Common patterns (ETL, time series, multi-file processing)
78
-
79
- **Quick Example**:
80
- ```python
81
- import dask.dataframe as dd
82
-
83
- # Read multiple files as single DataFrame
84
- ddf = dd.read_csv('data/2024-*.csv')
85
-
86
- # Operations are lazy until compute()
87
- filtered = ddf[ddf['value'] > 100]
88
- result = filtered.groupby('category').mean().compute()
89
- ```
90
-
91
- **Key Points**:
92
- - Operations are lazy (build task graph) until `.compute()` called
93
- - Use `map_partitions` for efficient custom operations
94
- - Convert to DataFrame early when working with structured data from other sources
95
-
96
- ### 2. Arrays - Parallel NumPy Operations
97
-
98
- **Purpose**: Extend NumPy capabilities to datasets larger than memory using blocked algorithms.
99
-
100
- **When to Use**:
101
- - Arrays exceed available RAM
102
- - NumPy operations need parallelization
103
- - Working with scientific datasets (HDF5, Zarr, NetCDF)
104
- - Need parallel linear algebra or array operations
105
-
106
- **Reference Documentation**: For comprehensive guidance on Dask Arrays, refer to `references/arrays.md` which includes:
107
- - Creating arrays (from NumPy, random, from disk)
108
- - Chunking strategies and optimization
109
- - Common operations (arithmetic, reductions, linear algebra)
110
- - Custom operations with `map_blocks`
111
- - Integration with HDF5, Zarr, and XArray
112
-
113
- **Quick Example**:
114
- ```python
115
- import dask.array as da
116
-
117
- # Create large array with chunks
118
- x = da.random.random((100000, 100000), chunks=(10000, 10000))
119
-
120
- # Operations are lazy
121
- y = x + 100
122
- z = y.mean(axis=0)
123
-
124
- # Compute result
125
- result = z.compute()
126
- ```
127
-
128
- **Key Points**:
129
- - Chunk size is critical (aim for ~100 MB per chunk)
130
- - Operations work on chunks in parallel
131
- - Rechunk data when needed for efficient operations
132
- - Use `map_blocks` for operations not available in Dask
133
-
134
- ### 3. Bags - Parallel Processing of Unstructured Data
135
-
136
- **Purpose**: Process unstructured or semi-structured data (text, JSON, logs) with functional operations.
137
-
138
- **When to Use**:
139
- - Processing text files, logs, or JSON records
140
- - Data cleaning and ETL before structured analysis
141
- - Working with Python objects that don't fit array/dataframe formats
142
- - Need memory-efficient streaming processing
143
-
144
- **Reference Documentation**: For comprehensive guidance on Dask Bags, refer to `references/bags.md` which includes:
145
- - Reading text and JSON files
146
- - Functional operations (map, filter, fold, groupby)
147
- - Converting to DataFrames
148
- - Common patterns (log analysis, JSON processing, text processing)
149
- - Performance considerations
150
-
151
- **Quick Example**:
152
- ```python
153
- import dask.bag as db
154
- import json
155
-
156
- # Read and parse JSON files
157
- bag = db.read_text('logs/*.json').map(json.loads)
158
-
159
- # Filter and transform
160
- valid = bag.filter(lambda x: x['status'] == 'valid')
161
- processed = valid.map(lambda x: {'id': x['id'], 'value': x['value']})
162
-
163
- # Convert to DataFrame for analysis
164
- ddf = processed.to_dataframe()
165
- ```
166
-
167
- **Key Points**:
168
- - Use for initial data cleaning, then convert to DataFrame/Array
169
- - Use `foldby` instead of `groupby` for better performance
170
- - Operations are streaming and memory-efficient
171
- - Convert to structured formats (DataFrame) for complex operations
172
-
173
- ### 4. Futures - Task-Based Parallelization
174
-
175
- **Purpose**: Build custom parallel workflows with fine-grained control over task execution and dependencies.
176
-
177
- **When to Use**:
178
- - Building dynamic, evolving workflows
179
- - Need immediate task execution (not lazy)
180
- - Computations depend on runtime conditions
181
- - Implementing custom parallel algorithms
182
- - Need stateful computations
183
-
184
- **Reference Documentation**: For comprehensive guidance on Dask Futures, refer to `references/futures.md` which includes:
185
- - Setting up distributed client
186
- - Submitting tasks and working with futures
187
- - Task dependencies and data movement
188
- - Advanced coordination (queues, locks, events, actors)
189
- - Common patterns (parameter sweeps, dynamic tasks, iterative algorithms)
190
-
191
- **Quick Example**:
192
- ```python
193
- from dask.distributed import Client
194
-
195
- client = Client() # Create local cluster
196
-
197
- # Submit tasks (executes immediately)
198
- def process(x):
199
- return x ** 2
200
-
201
- futures = client.map(process, range(100))
202
-
203
- # Gather results
204
- results = client.gather(futures)
205
-
206
- client.close()
207
- ```
208
-
209
- **Key Points**:
210
- - Requires distributed client (even for single machine)
211
- - Tasks execute immediately when submitted
212
- - Pre-scatter large data to avoid repeated transfers
213
- - ~1ms overhead per task (not suitable for millions of tiny tasks)
214
- - Use actors for stateful workflows
215
-
216
- ### 5. Schedulers - Execution Backends
217
-
218
- **Purpose**: Control how and where Dask tasks execute (threads, processes, distributed).
219
-
220
- **When to Choose Scheduler**:
221
- - **Threads** (default): NumPy/Pandas operations, GIL-releasing libraries, shared memory benefit
222
- - **Processes**: Pure Python code, text processing, GIL-bound operations
223
- - **Synchronous**: Debugging with pdb, profiling, understanding errors
224
- - **Distributed**: Need dashboard, multi-machine clusters, advanced features
225
-
226
- **Reference Documentation**: For comprehensive guidance on Dask Schedulers, refer to `references/schedulers.md` which includes:
227
- - Detailed scheduler descriptions and characteristics
228
- - Configuration methods (global, context manager, per-compute)
229
- - Performance considerations and overhead
230
- - Common patterns and troubleshooting
231
- - Thread configuration for optimal performance
232
-
233
- **Quick Example**:
234
- ```python
235
- import dask
236
- import dask.dataframe as dd
237
-
238
- # Use threads for DataFrame (default, good for numeric)
239
- ddf = dd.read_csv('data.csv')
240
- result1 = ddf.mean().compute() # Uses threads
241
-
242
- # Use processes for Python-heavy work
243
- import dask.bag as db
244
- bag = db.read_text('logs/*.txt')
245
- result2 = bag.map(python_function).compute(scheduler='processes')
246
-
247
- # Use synchronous for debugging
248
- dask.config.set(scheduler='synchronous')
249
- result3 = problematic_computation.compute() # Can use pdb
250
-
251
- # Use distributed for monitoring and scaling
252
- from dask.distributed import Client
253
- client = Client()
254
- result4 = computation.compute() # Uses distributed with dashboard
255
- ```
256
-
257
- **Key Points**:
258
- - Threads: Lowest overhead (~10 µs/task), best for numeric work
259
- - Processes: Avoids GIL (~10 ms/task), best for Python work
260
- - Distributed: Monitoring dashboard (~1 ms/task), scales to clusters
261
- - Can switch schedulers per computation or globally
262
-
263
- ## Best Practices
264
-
265
- For comprehensive performance optimization guidance, memory management strategies, and common pitfalls to avoid, refer to `references/best-practices.md`. Key principles include:
266
-
267
- ### Start with Simpler Solutions
268
- Before using Dask, explore:
269
- - Better algorithms
270
- - Efficient file formats (Parquet instead of CSV)
271
- - Compiled code (Numba, Cython)
272
- - Data sampling
273
-
274
- ### Critical Performance Rules
275
-
276
- **1. Don't Load Data Locally Then Hand to Dask**
277
- ```python
278
- # Wrong: Loads all data in memory first
279
- import pandas as pd
280
- df = pd.read_csv('large.csv')
281
- ddf = dd.from_pandas(df, npartitions=10)
282
-
283
- # Correct: Let Dask handle loading
284
- import dask.dataframe as dd
285
- ddf = dd.read_csv('large.csv')
286
- ```
287
-
288
- **2. Avoid Repeated compute() Calls**
289
- ```python
290
- # Wrong: Each compute is separate
291
- for item in items:
292
- result = dask_computation(item).compute()
293
-
294
- # Correct: Single compute for all
295
- computations = [dask_computation(item) for item in items]
296
- results = dask.compute(*computations)
297
- ```
298
-
299
- **3. Don't Build Excessively Large Task Graphs**
300
- - Increase chunk sizes if millions of tasks
301
- - Use `map_partitions`/`map_blocks` to fuse operations
302
- - Check task graph size: `len(ddf.__dask_graph__())`
303
-
304
- **4. Choose Appropriate Chunk Sizes**
305
- - Target: ~100 MB per chunk (or 10 chunks per core in worker memory)
306
- - Too large: Memory overflow
307
- - Too small: Scheduling overhead
308
-
309
- **5. Use the Dashboard**
310
- ```python
311
- from dask.distributed import Client
312
- client = Client()
313
- print(client.dashboard_link) # Monitor performance, identify bottlenecks
314
- ```
315
-
316
- ## Common Workflow Patterns
317
-
318
- ### ETL Pipeline
319
- ```python
320
- import dask.dataframe as dd
321
-
322
- # Extract: Read data
323
- ddf = dd.read_csv('raw_data/*.csv')
324
-
325
- # Transform: Clean and process
326
- ddf = ddf[ddf['status'] == 'valid']
327
- ddf['amount'] = ddf['amount'].astype('float64')
328
- ddf = ddf.dropna(subset=['important_col'])
329
-
330
- # Load: Aggregate and save
331
- summary = ddf.groupby('category').agg({'amount': ['sum', 'mean']})
332
- summary.to_parquet('output/summary.parquet')
333
- ```
334
-
335
- ### Unstructured to Structured Pipeline
336
- ```python
337
- import dask.bag as db
338
- import json
339
-
340
- # Start with Bag for unstructured data
341
- bag = db.read_text('logs/*.json').map(json.loads)
342
- bag = bag.filter(lambda x: x['status'] == 'valid')
343
-
344
- # Convert to DataFrame for structured analysis
345
- ddf = bag.to_dataframe()
346
- result = ddf.groupby('category').mean().compute()
347
- ```
348
-
349
- ### Large-Scale Array Computation
350
- ```python
351
- import dask.array as da
352
-
353
- # Load or create large array
354
- x = da.from_zarr('large_dataset.zarr')
355
-
356
- # Process in chunks
357
- normalized = (x - x.mean()) / x.std()
358
-
359
- # Save result (use mode= for overwrite; zarr_array_kwargs for compression)
360
- da.to_zarr(normalized, 'normalized.zarr', mode='w')
361
- ```
362
-
363
- ### Custom Parallel Workflow
364
- ```python
365
- from dask.distributed import Client
366
-
367
- client = Client()
368
-
369
- # Scatter large dataset once
370
- data = client.scatter(large_dataset)
371
-
372
- # Process in parallel with dependencies
373
- futures = []
374
- for param in parameters:
375
- future = client.submit(process, data, param)
376
- futures.append(future)
377
-
378
- # Gather results
379
- results = client.gather(futures)
380
- ```
381
-
382
- ## Selecting the Right Component
383
-
384
- Use this decision guide to choose the appropriate Dask component:
385
-
386
- **Data Type**:
387
- - Tabular data → **DataFrames**
388
- - Numeric arrays → **Arrays**
389
- - Text/JSON/logs → **Bags** (then convert to DataFrame)
390
- - Custom Python objects → **Bags** or **Futures**
391
-
392
- **Operation Type**:
393
- - Standard pandas operations → **DataFrames**
394
- - Standard NumPy operations → **Arrays**
395
- - Custom parallel tasks → **Futures**
396
- - Text processing/ETL → **Bags**
397
-
398
- **Control Level**:
399
- - High-level, automatic → **DataFrames/Arrays**
400
- - Low-level, manual → **Futures**
401
-
402
- **Workflow Type**:
403
- - Static computation graph → **DataFrames/Arrays/Bags**
404
- - Dynamic, evolving → **Futures**
405
-
406
- ## Integration Considerations
407
-
408
- ### File Formats
409
- - **Efficient**: Parquet, HDF5, Zarr (columnar, compressed, parallel-friendly)
410
- - **Compatible but slower**: CSV (use for initial ingestion only)
411
- - **For Arrays**: HDF5, Zarr, NetCDF
412
-
413
- ### Conversion Between Collections
414
- ```python
415
- # Bag → DataFrame
416
- ddf = bag.to_dataframe()
417
-
418
- # DataFrame → Array (for numeric data)
419
- arr = ddf.to_dask_array(lengths=True)
420
-
421
- # Array → DataFrame
422
- ddf = dd.from_dask_array(arr, columns=['col1', 'col2'])
423
- ```
424
-
425
- ### With Other Libraries
426
- - **XArray**: Wraps Dask arrays with labeled dimensions (geospatial, imaging)
427
- - **Dask-ML**: Machine learning with scikit-learn compatible APIs
428
- - **Distributed**: Advanced cluster management and monitoring
429
-
430
- ## Debugging and Development
431
-
432
- ### Iterative Development Workflow
433
-
434
- 1. **Test on small data with synchronous scheduler**:
435
- ```python
436
- dask.config.set(scheduler='synchronous')
437
- result = computation.compute() # Can use pdb, easy debugging
438
- ```
439
-
440
- 2. **Validate with threads on sample**:
441
- ```python
442
- sample = ddf.head(1000) # Small sample
443
- # Test logic, then scale to full dataset
444
- ```
445
-
446
- 3. **Scale with distributed for monitoring**:
447
- ```python
448
- from dask.distributed import Client
449
- client = Client()
450
- print(client.dashboard_link) # Monitor performance
451
- result = computation.compute()
452
- ```
453
-
454
- ### Common Issues
455
-
456
- **Memory Errors**:
457
- - Decrease chunk sizes
458
- - Use `persist()` strategically and delete when done
459
- - Check for memory leaks in custom functions
460
-
461
- **Slow Start**:
462
- - Task graph too large (increase chunk sizes)
463
- - Use `map_partitions` or `map_blocks` to reduce tasks
464
-
465
- **Poor Parallelization**:
466
- - Chunks too large (increase number of partitions)
467
- - Using threads with Python code (switch to processes)
468
- - Data dependencies preventing parallelism
469
-
470
- ## Reference Files
471
-
472
- All reference documentation files can be read as needed for detailed information:
473
-
474
- - `references/dataframes.md` - Complete Dask DataFrame guide
475
- - `references/arrays.md` - Complete Dask Array guide
476
- - `references/bags.md` - Complete Dask Bag guide
477
- - `references/futures.md` - Complete Dask Futures and distributed computing guide
478
- - `references/schedulers.md` - Complete scheduler selection and configuration guide
479
- - `references/best-practices.md` - Comprehensive performance optimization and troubleshooting
480
-
481
- Load these files when users need detailed information about specific Dask components, operations, or patterns beyond the quick guidance provided here.
482
-