@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,408 +0,0 @@
1
- ---
2
- name: lamindb
3
- description: Use when working with LaminDB, the open-source lineage-native lakehouse for biological datasets and models. Covers setup, artifact registration, query/search, lineage tracking, validation, ontology-backed annotation with Bionty, collections, branches, storage, and workflow integrations.
4
- license: Apache-2.0 license
5
- metadata:
6
- version: "1.1"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # LaminDB
11
-
12
- ## Overview
13
-
14
- LaminDB is an open-source, lineage-native lakehouse for biology. It makes datasets and models queryable, traceable, validated, reproducible, and FAIR (Findable, Accessible, Interoperable, Reusable) while storing data in open formats across local filesystems, S3, GCS, Hugging Face, SQLite, and Postgres.
15
-
16
- **Core Value Proposition:**
17
- - **Queryability**: Search and filter artifacts, records, runs, features, schemas, and collections
18
- - **Traceability**: Track inputs, outputs, parameters, source code, and environments for notebooks, scripts, functions, and pipelines
19
- - **Validation**: Curate DataFrame, AnnData, SpatialData, TileDB-SOMA, Parquet, Zarr, and other biological formats with schemas
20
- - **FAIR Compliance**: Standardize annotations with Bionty-backed ontologies and custom registries
21
- - **Change management**: Organize work with projects, branches, spaces, collections, and saved notes or plans
22
-
23
- ## When to Use This Skill
24
-
25
- Use this skill when:
26
-
27
- - **Managing biological datasets**: scRNA-seq, bulk RNA-seq, spatial transcriptomics, flow cytometry, multi-modal data, EHR data
28
- - **Tracking computational workflows**: Notebooks, scripts, functions, shell scripts, and pipeline execution (Nextflow, Snakemake, Redun)
29
- - **Curating and validating data**: Schema validation, standardization, ontology-based annotation
30
- - **Working with biological ontologies**: Genes, proteins, cell types, tissues, diseases, pathways (via Bionty)
31
- - **Building data lakehouses**: Unified query interface across multiple datasets
32
- - **Ensuring reproducibility**: Automatic versioning, lineage tracking, environment capture
33
- - **Integrating ML pipelines**: Connecting with Weights & Biases, MLflow, Hugging Face, Lightning, scVI-tools
34
- - **Deploying data infrastructure**: Setting up local or cloud-based data management systems
35
- - **Collaborating on datasets**: Sharing curated, annotated data with standardized metadata
36
-
37
- ## Core Capabilities
38
-
39
- LaminDB provides six interconnected capability areas, each documented in detail in the references folder.
40
-
41
- ### 1. Core Concepts and Data Lineage
42
-
43
- **Core entities:**
44
- - **Artifacts**: Versioned datasets (DataFrame, AnnData, Parquet, Zarr, etc.)
45
- - **Records & ULabels**: Experimental entities, typed records, and simple labels
46
- - **Collections**: Versioned, immutable sets of artifacts
47
- - **Runs & Transforms**: Computational lineage tracking (what code produced what data)
48
- - **Features**: Typed metadata fields for annotation and querying
49
- - **Projects, Branches & Spaces**: Project grouping, change management, and access boundaries
50
-
51
- **Key workflows:**
52
- - Create and version artifacts from files or Python objects
53
- - Track notebook/script execution with `ln.track()` and `ln.finish()`
54
- - Track function workflows with `@ln.flow()` and `@ln.step()`
55
- - Annotate artifacts with records, ulabels, projects, and typed features
56
- - Visualize data lineage graphs with `artifact.view_lineage()`
57
- - Query by provenance (find all outputs from specific code/inputs)
58
-
59
- **Reference:** `references/core-concepts.md` - Read this for detailed information on artifacts, records, runs, transforms, features, versioning, and lineage tracking.
60
-
61
- ### 2. Data Management and Querying
62
-
63
- **Query capabilities:**
64
- - Registry exploration and lookup with auto-complete
65
- - Single record retrieval with `get()`, `one()`, `one_or_none()`
66
- - Filtering with comparison operators (`__gt`, `__lte`, `__contains`, `__startswith`)
67
- - Feature-based queries, including expression-style queries with `Feature` objects
68
- - Cross-registry traversal with double-underscore syntax
69
- - Full-text search across registries
70
- - Advanced logical queries with `ln.Q` objects (AND, OR, NOT)
71
- - Streaming large datasets without loading into memory
72
-
73
- **Key workflows:**
74
- - Browse artifacts with filters and ordering
75
- - Query by features, creation date, creator, size, etc.
76
- - Stream large files in chunks or with array slicing
77
- - Organize data with hierarchical keys
78
- - Group artifacts into collections
79
-
80
- **Reference:** `references/data-management.md` - Read this for comprehensive query patterns, filtering examples, streaming strategies, and data organization best practices.
81
-
82
- ### 3. Annotation and Validation
83
-
84
- **Curation process:**
85
- 1. **Validation**: Confirm datasets match desired schemas
86
- 2. **Standardization**: Fix typos, map synonyms to canonical terms
87
- 3. **Annotation**: Link datasets to metadata entities for queryability
88
-
89
- **Schema types:**
90
- - **Flexible schemas**: Validate only known columns, allow additional metadata
91
- - **Minimal required schemas**: Specify essential columns, permit extras
92
- - **Strict schemas**: Complete control over structure and values
93
-
94
- **Supported data types:**
95
- - DataFrames (Parquet, CSV)
96
- - AnnData (single-cell genomics)
97
- - MuData (multi-modal)
98
- - SpatialData (spatial transcriptomics)
99
- - TileDB-SOMA (scalable arrays)
100
-
101
- **Key workflows:**
102
- - Define features and schemas for data validation
103
- - Use `DataFrameCurator`, `AnnDataCurator`, `SpatialDataCurator`, or `TiledbsomaExperimentCurator` for validation
104
- - Standardize values with `.cat.standardize()`
105
- - Map to ontologies with `.cat.add_ontology()`
106
- - Save curated artifacts with schema linkage
107
- - Query validated datasets by features
108
-
109
- **Reference:** `references/annotation-validation.md` - Read this for detailed curation workflows, schema design patterns, handling validation errors, and best practices.
110
-
111
- ### 4. Biological Ontologies
112
-
113
- **Available ontologies (via Bionty):**
114
- - Genes (Ensembl), Proteins (UniProt)
115
- - Cell types (CL), Cell lines (CLO)
116
- - Tissues (Uberon), Diseases (Mondo, DOID)
117
- - Phenotypes (HPO), Pathways (GO)
118
- - Experimental factors (EFO), Developmental stages
119
- - Organisms (NCBItaxon), Drugs (DrugBank)
120
-
121
- **Key workflows:**
122
- - Import public ontologies with `bt.CellType.import_source()`
123
- - Search ontologies with keyword or exact matching
124
- - Standardize terms using synonym mapping
125
- - Explore hierarchical relationships (parents, children, ancestors)
126
- - Validate data against ontology terms
127
- - Annotate datasets with ontology records
128
- - Create custom terms and hierarchies
129
- - Handle multi-organism contexts (human, mouse, etc.)
130
-
131
- **Reference:** `references/ontologies.md` - Read this for comprehensive ontology operations, standardization strategies, hierarchy navigation, and annotation workflows.
132
-
133
- ### 5. Integrations
134
-
135
- **Workflow managers:**
136
- - Nextflow: Track pipeline processes and outputs
137
- - Snakemake: Integrate into Snakemake rules
138
- - Redun: Combine with Redun task tracking
139
- - Lightning: Persist checkpoints and training metadata
140
-
141
- **MLOps platforms:**
142
- - Weights & Biases: Link experiments with data artifacts
143
- - MLflow: Track models and experiments
144
- - Hugging Face: Track model fine-tuning
145
- - scVI-tools: Single-cell analysis workflows
146
-
147
- **Storage systems:**
148
- - Local filesystem, AWS S3, Google Cloud Storage
149
- - S3-compatible (MinIO, Cloudflare R2)
150
- - HTTP/HTTPS endpoints (read-only)
151
- - HuggingFace datasets
152
-
153
- **Array stores:**
154
- - TileDB-SOMA (with cellxgene support)
155
- - DuckDB for SQL queries on Parquet files
156
-
157
- **Visualization:**
158
- - Vitessce for interactive spatial/single-cell visualization
159
-
160
- **Version control:**
161
- - Git integration for source code tracking
162
-
163
- **Reference:** `references/integrations.md` - Read this for integration patterns, code examples, and troubleshooting for third-party systems.
164
-
165
- ### 6. Setup and Deployment
166
-
167
- **Installation:**
168
- - Current stable baseline: `lamindb==2.5.1` (released 2026-06-01; Python >=3.10, <=3.14)
169
- - Basic: `uv pip install 'lamindb==2.5.1'`
170
- - With extras: `uv pip install 'lamindb[gcp,zarr-v2,fcs]==2.5.1'`
171
- - Minimal namespace only: `uv pip install 'lamindb-core==2.5.1'`
172
- - Bionty module: included in the LaminDB docs and available as `uv pip install 'bionty==2.4.0'`
173
- - Optional modules: pin reviewed releases for wetlab or clinical schema modules rather than installing floating latest versions
174
-
175
- **Instance types:**
176
- - Local SQLite (development)
177
- - Cloud storage + SQLite (small teams)
178
- - Cloud storage + PostgreSQL (production)
179
-
180
- **Storage options:**
181
- - Local filesystem
182
- - AWS S3 with configurable regions and permissions
183
- - Google Cloud Storage
184
- - S3-compatible endpoints (MinIO, Cloudflare R2)
185
-
186
- **Configuration:**
187
- - Cache management for cloud files
188
- - Multi-user system configurations
189
- - Git repository sync
190
- - Named environment variables for credentials and connection URLs
191
-
192
- **Deployment patterns:**
193
- - Local dev → Cloud production migration
194
- - Multi-region deployments
195
- - Shared storage with personal instances
196
-
197
- **Reference:** `references/setup-deployment.md` - Read this for detailed installation, configuration, storage setup, database management, security best practices, and troubleshooting.
198
-
199
- ## Safety and Security Defaults
200
-
201
- When helping with LaminDB setup or integrations:
202
-
203
- - Never display, log, or transmit actual API keys, cloud credentials, database passwords, or full connection strings that include secrets.
204
- - Prefer IAM roles, workload identity, secret managers, or named environment variables such as `LAMIN_DB_URL`, `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, and `GOOGLE_APPLICATION_CREDENTIALS`; only check whether a named variable is present, not its value.
205
- - Before saving content from REST APIs, external databases, or user-provided files, validate and sanitize it with an explicit schema or curator.
206
- - For reproducible installs, pin package versions or use a lock file. Floating installs are acceptable only when the user explicitly wants the latest upstream release.
207
-
208
- ## Common Use Case Workflows
209
-
210
- ### Use Case 1: Single-Cell RNA-seq Analysis with Ontology Validation
211
-
212
- ```python
213
- import lamindb as ln
214
- import bionty as bt
215
- import anndata as ad
216
-
217
- # Start tracking a notebook/script run
218
- ln.track(params={"analysis": "scRNA-seq QC and annotation"})
219
-
220
- # Import cell type ontology
221
- bt.CellType.import_source()
222
-
223
- # Load data
224
- adata = ad.read_h5ad("raw_counts.h5ad")
225
-
226
- # Validate and standardize cell types
227
- adata.obs["cell_type"] = bt.CellType.standardize(adata.obs["cell_type"])
228
-
229
- # Curate with schema
230
- curator = ln.curators.AnnDataCurator(adata, schema)
231
- curator.validate()
232
- artifact = curator.save_artifact(key="scrna/validated.h5ad")
233
-
234
- # Link ontology-backed annotations for queryability
235
- cell_types = bt.CellType.from_values(adata.obs["cell_type"])
236
- artifact.cell_types.add(*cell_types)
237
-
238
- ln.finish()
239
- ```
240
-
241
- ### Use Case 2: Building a Queryable Data Lakehouse
242
-
243
- ```python
244
- import lamindb as ln
245
-
246
- # Register multiple experiments
247
- for i, file in enumerate(data_files):
248
- artifact = ln.Artifact.from_anndata(
249
- ad.read_h5ad(file),
250
- key=f"scrna/batch_{i}.h5ad",
251
- description=f"scRNA-seq batch {i}"
252
- ).save()
253
-
254
- # Annotate with features
255
- artifact.features.set_values({
256
- "batch": i,
257
- "tissue": tissues[i],
258
- "condition": conditions[i]
259
- })
260
-
261
- # Query across all experiments by annotated features
262
- immune_datasets = ln.Artifact.filter(
263
- key__startswith="scrna/",
264
- tissue="PBMC",
265
- condition="treated"
266
- ).to_dataframe()
267
-
268
- # Load specific datasets
269
- for artifact in immune_datasets:
270
- adata = artifact.load()
271
- # Analyze
272
- ```
273
-
274
- ### Use Case 3: ML Pipeline with W&B Integration
275
-
276
- ```python
277
- import lamindb as ln
278
- import wandb
279
-
280
- # Initialize both systems
281
- wandb.init(project="drug-response", name="exp-42")
282
- ln.track(params={"model": "random_forest", "n_estimators": 100})
283
-
284
- # Load training data from LaminDB
285
- train_artifact = ln.Artifact.get(key="datasets/train.parquet")
286
- train_data = train_artifact.load()
287
-
288
- # Train model
289
- model = train_model(train_data)
290
-
291
- # Log to W&B
292
- wandb.log({"accuracy": 0.95})
293
-
294
- # Save model in LaminDB with W&B linkage
295
- import joblib
296
- joblib.dump(model, "model.pkl")
297
- model_artifact = ln.Artifact("model.pkl", key="models/exp-42.pkl").save()
298
- model_artifact.features.set_values({"wandb_run_id": wandb.run.id})
299
-
300
- ln.finish()
301
- wandb.finish()
302
- ```
303
-
304
- ### Use Case 4: Nextflow Pipeline Integration
305
-
306
- ```python
307
- # In Nextflow process script
308
- import lamindb as ln
309
-
310
- ln.track()
311
-
312
- # Load input artifact
313
- input_artifact = ln.Artifact.get(key="raw/batch_${batch_id}.fastq.gz")
314
- input_path = input_artifact.cache()
315
-
316
- # Process (alignment, quantification, etc.)
317
- # ... Nextflow process logic ...
318
-
319
- # Save output
320
- output_artifact = ln.Artifact(
321
- "counts.csv",
322
- key="processed/batch_${batch_id}_counts.csv"
323
- ).save()
324
-
325
- ln.finish()
326
- ```
327
-
328
- For native Nextflow projects, prefer the `nf-lamin` plugin and current `nextflow.config` patterns when available; use inline Python tracking for small or custom pipeline steps.
329
-
330
- ## Getting Started Checklist
331
-
332
- To start using LaminDB effectively:
333
-
334
- 1. **Installation & Setup** (`references/setup-deployment.md`)
335
- - Install pinned LaminDB and required extras
336
- - Authenticate with `lamin login`
337
- - Initialize instance with `lamin init --storage ...`
338
-
339
- 2. **Learn Core Concepts** (`references/core-concepts.md`)
340
- - Understand Artifacts, Records, Runs, Transforms
341
- - Practice creating and retrieving artifacts
342
- - Implement `ln.track()`/`ln.finish()` or `@ln.flow()`/`@ln.step()` in workflows
343
-
344
- 3. **Master Querying** (`references/data-management.md`)
345
- - Practice filtering and searching registries
346
- - Learn feature-based queries and expression-style filters
347
- - Experiment with streaming large files
348
-
349
- 4. **Set Up Validation** (`references/annotation-validation.md`)
350
- - Define features relevant to research domain
351
- - Create schemas for data types
352
- - Practice curation workflows
353
-
354
- 5. **Integrate Ontologies** (`references/ontologies.md`)
355
- - Import relevant biological ontologies (genes, cell types, etc.)
356
- - Validate existing annotations
357
- - Standardize metadata with ontology terms
358
-
359
- 6. **Connect Tools** (`references/integrations.md`)
360
- - Integrate with existing workflow managers
361
- - Link ML platforms for experiment tracking
362
- - Configure cloud storage and compute
363
-
364
- ## Key Principles
365
-
366
- Follow these principles when working with LaminDB:
367
-
368
- 1. **Track everything**: Use `ln.track()` at the start of every analysis for automatic lineage capture
369
-
370
- 2. **Validate early**: Define schemas and validate data before extensive analysis
371
-
372
- 3. **Use ontologies**: Leverage public biological ontologies for standardized annotations
373
-
374
- 4. **Organize with keys**: Structure artifact keys hierarchically (e.g., `project/experiment/batch/file.h5ad`)
375
-
376
- 5. **Query metadata first**: Filter and search before loading large files
377
-
378
- 6. **Version, don't duplicate**: Use built-in versioning instead of creating new keys for modifications
379
-
380
- 7. **Annotate with features**: Define typed features and use `artifact.features.set_values()` for queryable metadata
381
-
382
- 8. **Document thoroughly**: Add descriptions to artifacts, schemas, and transforms
383
-
384
- 9. **Leverage lineage**: Use `view_lineage()` to understand data provenance
385
-
386
- 10. **Start local, scale cloud**: Develop locally with SQLite, deploy to cloud with PostgreSQL
387
-
388
- ## Reference Files
389
-
390
- This skill includes comprehensive reference documentation organized by capability:
391
-
392
- - **`references/core-concepts.md`** - Artifacts, records, runs, transforms, features, versioning, lineage
393
- - **`references/data-management.md`** - Querying, filtering, searching, streaming, organizing data
394
- - **`references/annotation-validation.md`** - Schema design, curation workflows, validation strategies
395
- - **`references/ontologies.md`** - Biological ontology management, standardization, hierarchies
396
- - **`references/integrations.md`** - Workflow managers, MLOps platforms, storage systems, tools
397
- - **`references/setup-deployment.md`** - Installation, configuration, deployment, troubleshooting
398
-
399
- Read the relevant reference file(s) based on the specific LaminDB capability needed for the task at hand.
400
-
401
- ## Additional Resources
402
-
403
- - **Official Documentation**: https://docs.lamin.ai
404
- - **API Reference**: https://docs.lamin.ai/api
405
- - **GitHub Repository**: https://github.com/laminlabs/lamindb
406
- - **Tutorial**: https://docs.lamin.ai/tutorial
407
- - **FAQ**: https://docs.lamin.ai/faq
408
-
@@ -1,227 +0,0 @@
1
- ---
2
- name: latchbio-integration
3
- description: Build, register, debug, and operate bioinformatics workflows on Latch using the Python SDK, CLI, Latch Data and Registry, Nextflow, Snakemake, programmatic execution, and Latch MCP. Use when authoring or deploying Latch workflows, configuring resources or interfaces, moving data, integrating Registry, or launching and monitoring runs.
4
- license: MIT
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires network access and a Latch account. The current stable SDK requires Python 3.9+; Python 3.12 is recommended. Uses uv for installation. Docker is needed for local image builds, while remote registration is the CLI default.
7
- metadata:
8
- version: "2.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # LatchBio Integration
13
-
14
- ## Current Baseline
15
-
16
- This skill targets **Latch SDK 2.76.8**, released July 10, 2026. The package
17
- metadata supports Python 3.9–3.12 and declares Python 3.9+.
18
-
19
- Treat the installed package and its changelog as authoritative when a guide
20
- disagrees with the SDK. Some Latch guides retain older Python ranges or
21
- compatibility-specific pre-release pins, especially the Snakemake v2 tutorial.
22
- Never combine commands or imports from different tracks without checking their
23
- version requirements.
24
-
25
- ## When to Use
26
-
27
- Use this skill to:
28
-
29
- - Create or maintain Python SDK workflows and task graphs
30
- - Package and register Python, Nextflow, or Snakemake pipelines
31
- - Configure task CPU, memory, storage, GPU, caching, retries, and timeouts
32
- - Work with Latch Data through `LPath`, `LatchFile`, `LatchDir`, or the CLI
33
- - Read or update Latch Registry projects, tables, and records
34
- - Design workflow forms, launch plans, samplesheets, messages, and result links
35
- - Stage and debug workflow images with `latch register --staging` and `latch develop`
36
- - Launch and monitor workflows through Python or Latch MCP
37
- - Discover and use ready-to-run Latch workflows
38
-
39
- ## Route to the Right Reference
40
-
41
- Read only the references needed for the task:
42
-
43
- | Need | Reference |
44
- |---|---|
45
- | Python workflows, tasks, maps, conditions, caching | `references/workflow-creation.md` |
46
- | `LPath`, legacy file types, Latch URLs, data CLI | `references/data-management.md` |
47
- | Registry reads, transactions, samplesheets | `references/registry.md` |
48
- | CPU, memory, storage, GPU, dynamic resources | `references/resource-configuration.md` |
49
- | Nextflow and Snakemake packaging | `references/nextflow-snakemake.md` |
50
- | Metadata, forms, launch plans, messages, automations | `references/ui-and-automation.md` |
51
- | Registration, development, execution, monitoring | `references/operations-and-debugging.md` |
52
- | Ready-to-use workflows and `latch.verified` | `references/verified-workflows.md` |
53
- | Remote MCP setup and tool workflow | `references/latch-mcp.md` |
54
-
55
- Before relying on a symbol, run `scripts/inspect_latch_sdk.py` against the
56
- target SDK version. It performs local imports only and does not authenticate or
57
- make network requests.
58
-
59
- ## Installation and Authentication
60
-
61
- For a reproducible environment:
62
-
63
- ```bash
64
- uv venv --python 3.12
65
- source .venv/bin/activate
66
- uv pip install "latch==2.76.8"
67
- ```
68
-
69
- On Windows, use WSL for the documented Linux workflow tooling.
70
-
71
- Authenticate through the supported OAuth flow; do not read, print, copy, or
72
- parse `~/.latch/token` manually:
73
-
74
- ```bash
75
- latch login
76
- latch workspace
77
- ```
78
-
79
- Select a workspace non-interactively when its numeric ID is already known:
80
-
81
- ```bash
82
- latch workspace --id 12345
83
- ```
84
-
85
- `latch login` credentials are for the SDK and CLI. Latch MCP uses a separate
86
- OAuth authorization and its credentials cannot be reused for general SDK
87
- access.
88
-
89
- ## Fast Path
90
-
91
- Create and remotely register the maintained subprocess template:
92
-
93
- ```bash
94
- latch init covid-wf --template subprocess
95
- latch register --yes --open covid-wf
96
- ```
97
-
98
- Remote image building is the default. Use `--no-remote` only when a local
99
- Docker daemon is available and a local build is intentional.
100
-
101
- ## Minimal Python Workflow
102
-
103
- Keep workflow bodies declarative: invoke tasks and return their promises.
104
- Perform computation and side effects inside tasks.
105
-
106
- ```python
107
- from latch import small_task, workflow
108
-
109
-
110
- @small_task
111
- def reverse_complement(sequence: str) -> str:
112
- table = str.maketrans("ACGTacgt", "TGCAtgca")
113
- return sequence.translate(table)[::-1]
114
-
115
-
116
- @workflow
117
- def reverse_complement_workflow(sequence: str) -> str:
118
- """Return the reverse complement of a DNA sequence."""
119
- return reverse_complement(sequence=sequence)
120
- ```
121
-
122
- Use `@workflow(metadata)` when the generated interface needs custom labels,
123
- sections, validation rules, samplesheets, or documentation links. Use `LatchFile` or
124
- `LatchDir` for automatic task input staging and output upload; use `LPath` for
125
- imperative remote path operations.
126
-
127
- ## Recommended Development Lifecycle
128
-
129
- 1. **Inspect compatibility**
130
- - Confirm the installed SDK and Python version.
131
- - Identify whether the project is Python, Nextflow, the legacy Snakemake
132
- flag path, or the separately pinned Snakemake v2 tutorial track.
133
-
134
- 2. **Define a typed interface**
135
- - Annotate every workflow and task input and output.
136
- - Keep module import time free of network calls, data mutations, and secret
137
- retrieval. Isolate documented exceptions such as `workflow_reference`,
138
- which resolves the active workspace when its decorator is evaluated.
139
- - Use dataclasses and enums for structured parameters.
140
-
141
- 3. **Configure metadata and resources**
142
- - Match metadata parameter keys to the workflow signature.
143
- - Start with named task decorators, then use `custom_task` only when measured
144
- requirements justify it.
145
-
146
- 4. **Validate in the execution image**
147
-
148
- Fresh Nextflow and Snakemake projects must generate their
149
- version-compatible Python entrypoint before staging. In SDK 2.76.8, the
150
- staging branch does not generate one from `--nf-script` or `--snakefile`.
151
-
152
- ```bash
153
- latch register --staging .
154
- latch develop .
155
- ```
156
-
157
- Re-run staging registration after changing the Dockerfile or dependencies.
158
- Edits made inside the development container are not synced back.
159
-
160
- 5. **Register deliberately**
161
-
162
- ```bash
163
- latch register --yes --open .
164
- ```
165
-
166
- Useful controls:
167
-
168
- ```bash
169
- latch register --workspace-id 12345 .
170
- latch register --mark-as-release .
171
- latch register --workflow-module wf.custom_entrypoint .
172
- ```
173
-
174
- Duplicate registration exits with status `2`; it is not the same as a build
175
- failure.
176
-
177
- 6. **Launch only after reviewing cost and parameters**
178
- - Prefer the Console or Latch MCP for interactive operation.
179
- - Prefer `latch_cli.services.launch.launch_v2` for Python automation.
180
- - Do not use the deprecated `latch launch` CLI as a new integration pattern.
181
-
182
- 7. **Monitor and verify**
183
- - Check terminal status, task logs, result links, and scientific outputs.
184
- - Treat successful orchestration as necessary but not sufficient scientific
185
- validation.
186
-
187
- ## Operational Safety
188
-
189
- - Ask for confirmation before launching paid compute, especially GPU or large
190
- batch runs.
191
- - Ask for confirmation before `LPath.rmr`, `latch rmr`, Registry deletion, or
192
- overwriting shared destinations.
193
- - Never log secrets, SDK tokens, signed URLs, or secret values.
194
- - Call `get_secret()` only inside a task, use the returned value only for its
195
- intended service, and never return it as workflow output.
196
- - Do not pass untrusted strings through shell commands. Prefer argument lists
197
- with `subprocess.run(..., check=True)`.
198
- - Pin the SDK and workflow dependencies for releases. Upgrade only after
199
- reviewing the changelog and re-running staging tests.
200
- - Treat generated files as generated: customize the documented extension file
201
- rather than editing output that the CLI will overwrite.
202
-
203
- ## Inspect the Installed SDK
204
-
205
- From this skill directory:
206
-
207
- ```bash
208
- uv run --no-project --python 3.12 --with "latch==2.76.8" \
209
- python scripts/inspect_latch_sdk.py
210
- ```
211
-
212
- Use JSON output for automated comparisons:
213
-
214
- ```bash
215
- uv run --no-project --python 3.12 --with "latch==2.76.8" \
216
- python scripts/inspect_latch_sdk.py --json
217
- ```
218
-
219
- ## Authoritative Sources
220
-
221
- - Documentation index: https://wiki.latch.bio/llms.txt
222
- - Workflow and SDK guides: https://wiki.latch.bio/workflows/overview
223
- - SDK API reference: https://wiki.latch.bio/reference/sdk
224
- - PyPI package: https://pypi.org/project/latch/
225
- - SDK 2.76.8 release source: https://github.com/latchbio/latch/tree/0faa9dcd8186444ac008f50adf95d43f0fa30e06
226
- - SDK changelog: https://github.com/latchbio/latch/blob/0faa9dcd8186444ac008f50adf95d43f0fa30e06/CHANGELOG.md
227
- - Latch Console: https://console.latch.bio