@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,456 +0,0 @@
1
- ---
2
- name: tiledbvcf
3
- description: Efficient storage and retrieval of genomic variant data using TileDB. Scalable VCF/BCF ingestion, incremental sample addition, compressed storage, parallel queries, and export capabilities for population genomics.
4
- license: MIT license
5
- metadata:
6
- version: "1.1"
7
- skill-author: Jeremy Leipzig
8
- ---
9
-
10
- # TileDB-VCF
11
-
12
- ## Overview
13
-
14
- TileDB-VCF is a high-performance C++ library with Python and CLI interfaces for efficient storage and retrieval of genomic variant-call data. Built on TileDB's sparse array technology, it enables scalable ingestion of VCF/BCF files, incremental sample addition without expensive merging operations, and efficient parallel queries of variant data stored locally or in the cloud.
15
-
16
- ## When to Use This Skill
17
-
18
- This skill should be used when:
19
- - Learning TileDB-VCF concepts and workflows
20
- - Prototyping genomics analyses and pipelines
21
- - Working with small-to-medium datasets (< 1000 samples)
22
- - Need incremental addition of new samples to existing datasets
23
- - Require efficient querying of specific genomic regions across many samples
24
- - Working with cloud-stored variant data (S3, Azure, GCS)
25
- - Need to export subsets of large VCF datasets
26
- - Building variant databases for cohort studies
27
- - Educational projects and method development
28
- - Performance is critical for variant data operations
29
-
30
- ## Quick Start
31
-
32
- ### Installation
33
-
34
- **Preferred Method: Conda/Mamba**
35
- ```bash
36
- # Enter the following two lines if you are on a M1 Mac
37
- CONDA_SUBDIR=osx-64
38
- conda config --env --set subdir osx-64
39
-
40
- # Create the conda environment
41
- conda create -n tiledb-vcf "python<3.10"
42
- conda activate tiledb-vcf
43
-
44
- # Mamba is a faster and more reliable alternative to conda
45
- conda install -c conda-forge mamba
46
-
47
- # Install TileDB-Py and TileDB-VCF, align with other useful libraries
48
- mamba install -y -c conda-forge -c bioconda -c tiledb tiledb-py tiledbvcf-py pandas pyarrow numpy
49
- ```
50
-
51
- **Alternative: Docker Images**
52
- ```bash
53
- docker pull tiledb/tiledbvcf-py # Python interface
54
- docker pull tiledb/tiledbvcf-cli # Command-line interface
55
- ```
56
-
57
- ### Basic Examples
58
-
59
- **Create and populate a dataset:**
60
- ```python
61
- import tiledbvcf
62
-
63
- # Create a new dataset
64
- ds = tiledbvcf.Dataset(uri="my_dataset", mode="w",
65
- cfg=tiledbvcf.ReadConfig(memory_budget=1024))
66
-
67
- # Ingest VCF files (must be single-sample with indexes)
68
- # Requirements:
69
- # - VCFs must be single-sample (not multi-sample)
70
- # - Must have indexes: .csi (bcftools) or .tbi (tabix)
71
- ds.ingest_samples(["sample1.vcf.gz", "sample2.vcf.gz"])
72
- ```
73
-
74
- **Query variant data:**
75
- ```python
76
- # Open existing dataset for reading
77
- ds = tiledbvcf.Dataset(uri="my_dataset", mode="r")
78
-
79
- # Query specific regions and samples
80
- df = ds.read(
81
- attrs=["sample_name", "pos_start", "pos_end", "alleles", "fmt_GT"],
82
- regions=["chr1:1000000-2000000", "chr2:500000-1500000"],
83
- samples=["sample1", "sample2", "sample3"]
84
- )
85
- print(df.head())
86
- ```
87
-
88
- **Export to VCF:**
89
- ```python
90
- import os
91
-
92
- # Export two VCF samples
93
- ds.export(
94
- regions=["chr21:8220186-8405573"],
95
- samples=["HG00101", "HG00097"],
96
- output_format="v",
97
- output_dir=os.path.expanduser("~"),
98
- )
99
- ```
100
-
101
- ## Core Capabilities
102
-
103
- ### 1. Dataset Creation and Ingestion
104
-
105
- Create TileDB-VCF datasets and incrementally ingest variant data from multiple VCF/BCF files. This is appropriate for building population genomics databases and cohort studies.
106
-
107
- **Requirements:**
108
- - **Single-sample VCFs only**: Multi-sample VCFs are not supported
109
- - **Index files required**: VCF/BCF files must have indexes (.csi or .tbi)
110
-
111
- **Common operations:**
112
- - Create new datasets with optimized array schemas
113
- - Ingest single or multiple VCF/BCF files in parallel
114
- - Add new samples incrementally without re-processing existing data
115
- - Configure memory usage and compression settings
116
- - Handle various VCF formats and INFO/FORMAT fields
117
- - Resume interrupted ingestion processes
118
- - Validate data integrity during ingestion
119
-
120
-
121
- ### 2. Efficient Querying and Filtering
122
-
123
- Query variant data with high performance across genomic regions, samples, and variant attributes. This is appropriate for association studies, variant discovery, and population analysis.
124
-
125
- **Common operations:**
126
- - Query specific genomic regions (single or multiple)
127
- - Filter by sample names or sample groups
128
- - Extract specific variant attributes (position, alleles, genotypes, quality)
129
- - Access INFO and FORMAT fields efficiently
130
- - Combine spatial and attribute-based filtering
131
- - Stream large query results
132
- - Perform aggregations across samples or regions
133
-
134
-
135
- ### 3. Data Export and Interoperability
136
-
137
- Export data in various formats for downstream analysis or integration with other genomics tools. This is appropriate for sharing datasets, creating analysis subsets, or feeding other pipelines.
138
-
139
- **Common operations:**
140
- - Export to standard VCF/BCF formats
141
- - Generate TSV files with selected fields
142
- - Create sample/region-specific subsets
143
- - Maintain data provenance and metadata
144
- - Lossless data export preserving all annotations
145
- - Compressed output formats
146
- - Streaming exports for large datasets
147
-
148
-
149
- ### 4. Population Genomics Workflows
150
-
151
- TileDB-VCF excels at large-scale population genomics analyses requiring efficient access to variant data across many samples and genomic regions.
152
-
153
- **Common workflows:**
154
- - Genome-wide association studies (GWAS) data preparation
155
- - Rare variant burden testing
156
- - Population stratification analysis
157
- - Allele frequency calculations across populations
158
- - Quality control across large cohorts
159
- - Variant annotation and filtering
160
- - Cross-population comparative analysis
161
-
162
-
163
- ## Key Concepts
164
-
165
- ### Array Schema and Data Model
166
-
167
- **TileDB-VCF Data Model:**
168
- - Variants stored as sparse arrays with genomic coordinates as dimensions
169
- - Samples stored as attributes allowing efficient sample-specific queries
170
- - INFO and FORMAT fields preserved with original data types
171
- - Automatic compression and chunking for optimal storage
172
-
173
- **Schema Configuration:**
174
- ```python
175
- # Custom schema with specific tile extents
176
- config = tiledbvcf.ReadConfig(
177
- memory_budget=2048, # MB
178
- region_partition=(0, 3095677412), # Full genome
179
- sample_partition=(0, 10000) # Up to 10k samples
180
- )
181
- ```
182
-
183
- ### Coordinate Systems and Regions
184
-
185
- **Critical:** TileDB-VCF uses **1-based genomic coordinates** following VCF standard:
186
- - Positions are 1-based (first base is position 1)
187
- - Ranges are inclusive on both ends
188
- - Region "chr1:1000-2000" includes positions 1000-2000 (1001 bases total)
189
-
190
- **Region specification formats:**
191
- ```python
192
- # Single region
193
- regions = ["chr1:1000000-2000000"]
194
-
195
- # Multiple regions
196
- regions = ["chr1:1000000-2000000", "chr2:500000-1500000"]
197
-
198
- # Whole chromosome
199
- regions = ["chr1"]
200
-
201
- # BED-style (0-based, half-open converted internally)
202
- regions = ["chr1:999999-2000000"] # Equivalent to 1-based chr1:1000000-2000000
203
- ```
204
-
205
- ### Memory Management
206
-
207
- **Performance considerations:**
208
- 1. **Set appropriate memory budget** based on available system memory
209
- 2. **Use streaming queries** for very large result sets
210
- 3. **Partition large ingestions** to avoid memory exhaustion
211
- 4. **Configure tile cache** for repeated region access
212
- 5. **Use parallel ingestion** for multiple files
213
- 6. **Optimize region queries** by combining nearby regions
214
-
215
- ### Cloud Storage Integration
216
-
217
- TileDB-VCF seamlessly works with cloud storage:
218
- ```python
219
- # S3 dataset
220
- ds = tiledbvcf.Dataset(uri="s3://bucket/dataset", mode="r")
221
-
222
- # Azure Blob Storage
223
- ds = tiledbvcf.Dataset(uri="azure://container/dataset", mode="r")
224
-
225
- # Google Cloud Storage
226
- ds = tiledbvcf.Dataset(uri="gcs://bucket/dataset", mode="r")
227
- ```
228
-
229
- ## Common Pitfalls
230
-
231
- 1. **Memory exhaustion during ingestion:** Use appropriate memory budget and batch processing for large VCF files
232
- 2. **Inefficient region queries:** Combine nearby regions instead of many separate queries
233
- 3. **Missing sample names:** Ensure sample names in VCF headers match query sample specifications
234
- 4. **Coordinate system confusion:** Remember TileDB-VCF uses 1-based coordinates like VCF standard
235
- 5. **Large result sets:** Use streaming or pagination for queries returning millions of variants
236
- 6. **Cloud permissions:** Ensure proper authentication for cloud storage access
237
- 7. **Concurrent access:** Multiple writers to the same dataset can cause corruption—use appropriate locking
238
-
239
- ## CLI Usage
240
-
241
- TileDB-VCF provides a command-line interface with the following subcommands:
242
-
243
- **Available Subcommands:**
244
- - `create` - Creates an empty TileDB-VCF dataset
245
- - `store` - Ingests samples into a TileDB-VCF dataset
246
- - `export` - Exports data from a TileDB-VCF dataset
247
- - `list` - Lists all sample names present in a TileDB-VCF dataset
248
- - `stat` - Prints high-level statistics about a TileDB-VCF dataset
249
- - `utils` - Utils for working with a TileDB-VCF dataset
250
- - `version` - Print the version information and exit
251
-
252
- ```bash
253
- # Create empty dataset
254
- tiledbvcf create --uri my_dataset
255
-
256
- # Ingest samples (requires single-sample VCFs with indexes)
257
- tiledbvcf store --uri my_dataset --samples sample1.vcf.gz,sample2.vcf.gz
258
-
259
- # Export data
260
- tiledbvcf export --uri my_dataset \
261
- --regions "chr1:1000000-2000000" \
262
- --sample-names "sample1,sample2"
263
-
264
- # List all samples
265
- tiledbvcf list --uri my_dataset
266
-
267
- # Show dataset statistics
268
- tiledbvcf stat --uri my_dataset
269
- ```
270
-
271
- ## Advanced Features
272
-
273
- ### Allele Frequency Analysis
274
- ```python
275
- # Calculate allele frequencies
276
- af_df = tiledbvcf.read_allele_frequency(
277
- uri="my_dataset",
278
- regions=["chr1:1000000-2000000"],
279
- samples=["sample1", "sample2", "sample3"]
280
- )
281
- ```
282
-
283
- ### Sample Quality Control
284
- ```python
285
- # Perform sample QC
286
- qc_results = tiledbvcf.sample_qc(
287
- uri="my_dataset",
288
- samples=["sample1", "sample2"]
289
- )
290
- ```
291
-
292
- ### Custom Configurations
293
- ```python
294
- # Advanced configuration
295
- config = tiledbvcf.ReadConfig(
296
- memory_budget=4096,
297
- tiledb_config={
298
- "sm.tile_cache_size": "1000000000",
299
- "vfs.s3.region": "us-east-1"
300
- }
301
- )
302
- ```
303
-
304
-
305
- ## Resources
306
-
307
- ## Getting Help
308
-
309
- ### Open Source TileDB-VCF Resources
310
-
311
- **Open Source Documentation:**
312
- - TileDB Academy: https://cloud.tiledb.com/academy/
313
- - Population Genomics Guide: https://cloud.tiledb.com/academy/structure/life-sciences/population-genomics/
314
- - TileDB-VCF GitHub: https://github.com/TileDB-Inc/TileDB-VCF
315
-
316
- ### TileDB-Cloud Resources
317
-
318
- **For Large-Scale/Production Genomics:**
319
- - TileDB-Cloud Platform: https://cloud.tiledb.com
320
- - TileDB Academy (All Documentation): https://cloud.tiledb.com/academy/
321
-
322
- **Getting Started:**
323
- - Free account signup: https://cloud.tiledb.com
324
- - Contact: sales@tiledb.com for enterprise needs
325
-
326
- ## Scaling to TileDB-Cloud
327
-
328
- When your genomics workloads outgrow single-node processing, TileDB-Cloud provides enterprise-scale capabilities for production genomics pipelines.
329
-
330
- **Note**: This section covers TileDB-Cloud capabilities based on available documentation. For complete API details and current functionality, consult the official TileDB-Cloud documentation and API reference.
331
-
332
- ### Setting Up TileDB-Cloud
333
-
334
- **1. Create Account and Get API Token**
335
- ```bash
336
- # Sign up at https://cloud.tiledb.com
337
- # Generate API token in your account settings
338
- ```
339
-
340
- **2. Install TileDB-Cloud Python Client**
341
- ```bash
342
- # Base installation
343
- uv pip install tiledb-cloud
344
-
345
- # With genomics-specific functionality
346
- uv pip install tiledb-cloud[life-sciences]
347
- ```
348
-
349
- **3. Configure Authentication**
350
- ```bash
351
- # Set environment variable with your API token
352
- export TILEDB_REST_TOKEN="your_api_token"
353
- ```
354
-
355
- ```python
356
- import tiledb.cloud
357
-
358
- # Authentication is automatic via TILEDB_REST_TOKEN
359
- # No explicit login required in code
360
- ```
361
-
362
- ### Migrating from Open Source to TileDB-Cloud
363
-
364
- **Large-Scale Ingestion**
365
- ```python
366
- # TileDB-Cloud: Distributed VCF ingestion
367
- import tiledb.cloud.vcf
368
-
369
- # Use specialized VCF ingestion module
370
- # Note: Exact API requires TileDB-Cloud documentation
371
- # This represents the available functionality structure
372
- tiledb.cloud.vcf.ingestion.ingest_vcf_dataset(
373
- source="s3://my-bucket/vcf-files/",
374
- output="tiledb://my-namespace/large-dataset",
375
- namespace="my-namespace",
376
- acn="my-s3-credentials",
377
- ingest_resources={"cpu": "16", "memory": "64Gi"}
378
- )
379
- ```
380
-
381
- **Distributed Query Processing**
382
- ```python
383
- # TileDB-Cloud: VCF querying across distributed storage
384
- import tiledb.cloud.vcf
385
- import tiledbvcf
386
-
387
- # Define the dataset URI
388
- dataset_uri = "tiledb://TileDB-Inc/gvcf-1kg-dragen-v376"
389
-
390
- # Get all samples from the dataset
391
- ds = tiledbvcf.Dataset(dataset_uri, tiledb_config=cfg)
392
- samples = ds.samples()
393
-
394
- # Define attributes and ranges to query on
395
- attrs = ["sample_name", "fmt_GT", "fmt_AD", "fmt_DP"]
396
- regions = ["chr13:32396898-32397044", "chr13:32398162-32400268"]
397
-
398
- # Perform the read, which is executed in a distributed fashion
399
- df = tiledb.cloud.vcf.read(
400
- dataset_uri=dataset_uri,
401
- regions=regions,
402
- samples=samples,
403
- attrs=attrs,
404
- namespace="my-namespace", # specifies which account to charge
405
- )
406
- df.to_pandas()
407
- ```
408
-
409
- ### Enterprise Features
410
-
411
- **Data Sharing and Collaboration**
412
- ```python
413
- # TileDB-Cloud provides enterprise data sharing capabilities
414
- # through namespace-based permissions and group management
415
-
416
- # Access shared datasets via TileDB-Cloud URIs
417
- dataset_uri = "tiledb://shared-namespace/population-study"
418
-
419
- # Collaborate through shared notebooks and compute resources
420
- # (Specific API requires TileDB-Cloud documentation)
421
- ```
422
-
423
- **Cost Optimization**
424
- - **Serverless Compute**: Pay only for actual compute time
425
- - **Auto-scaling**: Automatically scale up/down based on workload
426
- - **Spot Instances**: Use cost-optimized compute for batch jobs
427
- - **Data Tiering**: Automatic hot/cold storage management
428
-
429
- **Security and Compliance**
430
- - **End-to-end Encryption**: Data encrypted in transit and at rest
431
- - **Access Controls**: Fine-grained permissions and audit logs
432
- - **HIPAA/SOC2 Compliance**: Enterprise security standards
433
- - **VPC Support**: Deploy in private cloud environments
434
-
435
- ### When to Migrate Checklist
436
-
437
- ✅ **Migrate to TileDB-Cloud if you have:**
438
- - [ ] Datasets > 1000 samples
439
- - [ ] Need to process > 100GB of VCF data
440
- - [ ] Require distributed computing
441
- - [ ] Multiple team members need access
442
- - [ ] Need enterprise security/compliance
443
- - [ ] Want cost-optimized serverless compute
444
- - [ ] Require 24/7 production uptime
445
-
446
- ### Getting Started with TileDB-Cloud
447
-
448
- 1. **Start Free**: TileDB-Cloud offers free tier for evaluation
449
- 2. **Migration Support**: TileDB team provides migration assistance
450
- 3. **Training**: Access to genomics-specific tutorials and examples
451
- 4. **Professional Services**: Custom deployment and optimization
452
-
453
- **Next Steps:**
454
- - Visit https://cloud.tiledb.com to create account
455
- - Review documentation at https://cloud.tiledb.com/academy/
456
- - Contact sales@tiledb.com for enterprise needs