@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,328 +0,0 @@
1
- ---
2
- name: scvelo
3
- description: RNA velocity analysis with scVelo. Estimate cell state transitions from unspliced/spliced mRNA dynamics, infer trajectory directions, compute latent time, and identify driver genes in single-cell RNA-seq data. Complements Scanpy/scVI-tools for trajectory inference.
4
- license: BSD-3-Clause
5
- compatibility: Requires Python 3.10+ with scvelo, scanpy, and anndata. Verified against scvelo 0.3.4, whose dynamical model and pl.scatter need pandas<3 and whose stochastic estimator needs numpy<2; the deterministic estimator works on current releases.
6
- metadata:
7
- version: "1.2"
8
- skill-author: Kuan-lin Huang
9
- ---
10
-
11
- # scVelo — RNA Velocity Analysis
12
-
13
- ## Overview
14
-
15
- scVelo is the leading Python package for RNA velocity analysis in single-cell RNA-seq data. It infers cell state transitions by modeling the kinetics of mRNA splicing — using the ratio of unspliced (pre-mRNA) to spliced (mature mRNA) abundances to determine whether a gene is being upregulated or downregulated in each cell. This allows reconstruction of developmental trajectories and identification of cell fate decisions without requiring time-course data.
16
-
17
- **Installation:** `uv pip install scvelo`
18
-
19
- **Key resources:**
20
- - Documentation: https://scvelo.readthedocs.io/
21
- - GitHub: https://github.com/theislab/scvelo
22
- - Paper: Bergen et al. (2020) Nature Biotechnology. PMID: 32747759
23
-
24
- ## When to Use This Skill
25
-
26
- Use scVelo when:
27
-
28
- - **Trajectory inference from snapshot data**: Determine which direction cells are differentiating
29
- - **Cell fate prediction**: Identify progenitor cells and their downstream fates
30
- - **Driver gene identification**: Find genes whose dynamics best explain observed trajectories
31
- - **Developmental biology**: Model hematopoiesis, neurogenesis, epithelial-to-mesenchymal transitions
32
- - **Latent time estimation**: Order cells along a pseudotime derived from splicing dynamics
33
- - **Complement to Scanpy**: Add directional information to UMAP embeddings
34
-
35
- ## Prerequisites
36
-
37
- scVelo requires count matrices for both **unspliced** and **spliced** RNA. These are generated by:
38
- 1. **STARsolo** or **kallisto|bustools** with `lamanno` mode
39
- 2. **velocyto** CLI: `velocyto run10x` / `velocyto run`
40
- 3. **alevin-fry** / **simpleaf** with spliced/unspliced output
41
-
42
- Data is stored in an `AnnData` object with `layers["spliced"]` and `layers["unspliced"]`.
43
-
44
- ## Standard RNA Velocity Workflow
45
-
46
- ### 1. Setup and Data Loading
47
-
48
- ```python
49
- import scvelo as scv
50
- import scanpy as sc
51
- import numpy as np
52
- import matplotlib.pyplot as plt
53
-
54
- # Configure settings
55
- scv.settings.verbosity = 3 # Show computation steps
56
- scv.settings.presenter_view = True
57
- scv.settings.set_figure_params('scvelo')
58
-
59
- # Load data (AnnData with spliced/unspliced layers)
60
- # Option A: Load from loom (velocyto output)
61
- adata = scv.read("cellranger_output.loom", cache=True)
62
-
63
- # Option B: Merge velocyto loom with Scanpy-processed AnnData
64
- adata_processed = sc.read_h5ad("processed.h5ad") # Has UMAP, clusters
65
- adata_velocity = scv.read("velocyto.loom")
66
- adata = scv.utils.merge(adata_processed, adata_velocity)
67
-
68
- # Verify layers
69
- print(adata)
70
- # obs × var: N × G
71
- # layers: 'spliced', 'unspliced' (required)
72
- # obsm['X_umap'] (required for visualization)
73
- ```
74
-
75
- ### 2. Preprocessing
76
-
77
- ```python
78
- # Filter and normalize. As of scVelo 0.3, filter_and_normalize() only filters
79
- # genes and normalizes per cell -- it no longer takes n_top_genes and no longer
80
- # log-transforms, so the log step and HVG selection come from Scanpy.
81
- scv.pp.filter_and_normalize(
82
- adata,
83
- min_shared_counts=20 # Minimum counts in spliced+unspliced
84
- )
85
- sc.pp.log1p(adata)
86
- sc.pp.highly_variable_genes(adata, n_top_genes=2000, subset=True)
87
-
88
- # Compute first and second order moments (means and variances)
89
- # knn_connectivities must be computed first
90
- sc.pp.neighbors(adata, n_neighbors=30, n_pcs=30)
91
- scv.pp.moments(
92
- adata,
93
- n_pcs=30,
94
- n_neighbors=30
95
- )
96
- ```
97
-
98
- ### 3. Velocity Estimation — Stochastic Model
99
-
100
- The stochastic model is fast and suitable for exploratory analysis:
101
-
102
- ```python
103
- # Stochastic velocity (faster, less accurate)
104
- scv.tl.velocity(adata, mode='stochastic')
105
- scv.tl.velocity_graph(adata)
106
-
107
- # Visualize
108
- scv.pl.velocity_embedding_stream(
109
- adata,
110
- basis='umap',
111
- color='leiden',
112
- title="RNA Velocity (Stochastic)"
113
- )
114
- ```
115
-
116
- ### 4. Velocity Estimation — Dynamical Model (Recommended)
117
-
118
- The dynamical model fits the full splicing kinetics and is more accurate:
119
-
120
- ```python
121
- # Recover dynamics (computationally intensive; ~10-30 min for 10K cells)
122
- scv.tl.recover_dynamics(adata, n_jobs=4)
123
-
124
- # Compute velocity from dynamical model
125
- scv.tl.velocity(adata, mode='dynamical')
126
- scv.tl.velocity_graph(adata)
127
- ```
128
-
129
- ### 5. Latent Time
130
-
131
- The dynamical model enables computation of a shared latent time (pseudotime):
132
-
133
- ```python
134
- # Compute latent time
135
- scv.tl.latent_time(adata)
136
-
137
- # Visualize latent time on UMAP
138
- scv.pl.scatter(
139
- adata,
140
- color='latent_time',
141
- color_map='gnuplot',
142
- size=80,
143
- title='Latent time'
144
- )
145
-
146
- # Identify top genes ordered by latent time
147
- top_genes = adata.var['fit_likelihood'].sort_values(ascending=False).index[:300]
148
- scv.pl.heatmap(
149
- adata,
150
- var_names=top_genes,
151
- sortby='latent_time',
152
- col_color='leiden',
153
- n_convolve=100
154
- )
155
- ```
156
-
157
- ### 6. Driver Gene Analysis
158
-
159
- ```python
160
- # Identify genes with highest velocity fit
161
- scv.tl.rank_velocity_genes(adata, groupby='leiden', min_corr=0.3)
162
- df = scv.DataFrame(adata.uns['rank_velocity_genes']['names'])
163
- print(df.head(10))
164
-
165
- # Speed and coherence
166
- scv.tl.velocity_confidence(adata)
167
- scv.pl.scatter(
168
- adata,
169
- c=['velocity_length', 'velocity_confidence'],
170
- cmap='coolwarm',
171
- perc=[5, 95]
172
- )
173
-
174
- # Phase portraits for specific genes
175
- scv.pl.velocity(adata, ['Cpe', 'Gnao1', 'Ins2'],
176
- ncols=3, figsize=(16, 4))
177
- ```
178
-
179
- ### 7. Velocity Arrows and Pseudotime
180
-
181
- ```python
182
- # Arrow plot on UMAP
183
- scv.pl.velocity_embedding(
184
- adata,
185
- arrow_length=3,
186
- arrow_size=2,
187
- color='leiden',
188
- basis='umap'
189
- )
190
-
191
- # Stream plot (cleaner visualization)
192
- scv.pl.velocity_embedding_stream(
193
- adata,
194
- basis='umap',
195
- color='leiden',
196
- smooth=0.8,
197
- min_mass=4
198
- )
199
-
200
- # Velocity pseudotime (alternative to latent time)
201
- scv.tl.velocity_pseudotime(adata)
202
- scv.pl.scatter(adata, color='velocity_pseudotime', cmap='gnuplot')
203
- ```
204
-
205
- ### 8. PAGA Trajectory Graph
206
-
207
- ```python
208
- # PAGA graph with velocity-informed transitions
209
- scv.tl.paga(adata, groups='leiden')
210
- df = scv.get_df(adata, 'paga/transitions_confidence', precision=2).T
211
- df.style.background_gradient(cmap='Blues').format('{:.2g}')
212
-
213
- # Plot PAGA with velocity
214
- scv.pl.paga(
215
- adata,
216
- basis='umap',
217
- size=50,
218
- alpha=0.1,
219
- min_edge_width=2,
220
- node_size_scale=1.5
221
- )
222
- ```
223
-
224
- ## Complete Workflow Script
225
-
226
- ```python
227
- import scvelo as scv
228
- import scanpy as sc
229
-
230
- def run_rna_velocity(adata, n_top_genes=2000, mode='dynamical', n_jobs=4):
231
- """
232
- Complete RNA velocity workflow.
233
-
234
- Args:
235
- adata: AnnData with 'spliced' and 'unspliced' layers, UMAP in obsm
236
- n_top_genes: Number of top HVGs for velocity
237
- mode: 'stochastic' (fast) or 'dynamical' (accurate)
238
- n_jobs: Parallel jobs for dynamical model
239
-
240
- Returns:
241
- Processed AnnData with velocity information
242
- """
243
- scv.settings.verbosity = 2
244
-
245
- # 1. Preprocessing (scVelo 0.3 dropped log/HVG from filter_and_normalize)
246
- scv.pp.filter_and_normalize(adata, min_shared_counts=20)
247
- sc.pp.log1p(adata)
248
- sc.pp.highly_variable_genes(adata, n_top_genes=n_top_genes, subset=True)
249
-
250
- if 'neighbors' not in adata.uns:
251
- sc.pp.neighbors(adata, n_neighbors=30)
252
-
253
- scv.pp.moments(adata, n_pcs=30, n_neighbors=30)
254
-
255
- # 2. Velocity estimation
256
- if mode == 'dynamical':
257
- scv.tl.recover_dynamics(adata, n_jobs=n_jobs)
258
-
259
- scv.tl.velocity(adata, mode=mode)
260
- scv.tl.velocity_graph(adata)
261
-
262
- # 3. Downstream analyses
263
- if mode == 'dynamical':
264
- scv.tl.latent_time(adata)
265
- scv.tl.rank_velocity_genes(adata, groupby='leiden', min_corr=0.3)
266
-
267
- scv.tl.velocity_confidence(adata)
268
- scv.tl.velocity_pseudotime(adata)
269
-
270
- return adata
271
- ```
272
-
273
- ## Key Output Fields in AnnData
274
-
275
- After running the workflow, the following fields are added:
276
-
277
- | Location | Key | Description |
278
- |----------|-----|-------------|
279
- | `adata.layers` | `velocity` | RNA velocity per gene per cell |
280
- | `adata.layers` | `fit_t` | Fitted latent time per gene per cell |
281
- | `adata.obsm` | `velocity_umap` | 2D velocity vectors on UMAP |
282
- | `adata.obs` | `velocity_pseudotime` | Pseudotime from velocity |
283
- | `adata.obs` | `latent_time` | Latent time from dynamical model |
284
- | `adata.obs` | `velocity_length` | Speed of each cell |
285
- | `adata.obs` | `velocity_confidence` | Confidence score per cell |
286
- | `adata.var` | `fit_likelihood` | Gene-level model fit quality |
287
- | `adata.var` | `fit_alpha` | Transcription rate |
288
- | `adata.var` | `fit_beta` | Splicing rate |
289
- | `adata.var` | `fit_gamma` | Degradation rate |
290
- | `adata.uns` | `velocity_graph` | Cell-cell transition probability matrix |
291
-
292
- ## Velocity Models Comparison
293
-
294
- | Model | Speed | Accuracy | When to Use |
295
- |-------|-------|----------|-------------|
296
- | `stochastic` | Fast | Moderate | Exploratory; large datasets |
297
- | `deterministic` | Medium | Moderate | Simple linear kinetics |
298
- | `dynamical` | Slow | High | Publication-quality; identifies driver genes |
299
-
300
- ## Best Practices
301
-
302
- - **Start with stochastic mode** for exploration; switch to dynamical for final analysis
303
- - **Need good coverage of unspliced reads**: Short reads (< 100 bp) may miss intron coverage
304
- - **Minimum 2,000 cells**: RNA velocity is noisy with fewer cells
305
- - **Velocity should be coherent**: Arrows should follow known biology; randomness indicates issues
306
- - **k-NN bandwidth matters**: Too few neighbors → noisy velocity; too many → oversmoothed
307
- - **Sanity check**: Root cells (progenitors) should have high unspliced/spliced ratios for marker genes
308
- - **Dynamical model requires distinct kinetic states**: Works best for clear differentiation processes
309
-
310
- ## Troubleshooting
311
-
312
- | Problem | Solution |
313
- |---------|---------|
314
- | Missing unspliced layer | Re-run velocyto or use STARsolo with `--soloFeatures Gene Velocyto` |
315
- | Very few velocity genes | Lower `min_shared_counts`; check sequencing depth |
316
- | Random-looking arrows | Try different `n_neighbors` or velocity model |
317
- | Memory error with dynamical | Set `n_jobs=1`; reduce `n_top_genes` |
318
- | Negative velocity everywhere | Check that spliced/unspliced layers are not swapped |
319
-
320
- ## Additional Resources
321
-
322
- - **scVelo documentation**: https://scvelo.readthedocs.io/
323
- - **Tutorial notebooks**: https://scvelo.readthedocs.io/tutorials/
324
- - **GitHub**: https://github.com/theislab/scvelo
325
- - **Paper**: Bergen V et al. (2020) Nature Biotechnology. PMID: 32747759
326
- - **velocyto** (preprocessing): http://velocyto.org/
327
- - **CellRank** (fate prediction, extends scVelo): https://cellrank.readthedocs.io/
328
- - **dynamo** (metabolic labeling alternative): https://dynamo-release.readthedocs.io/
@@ -1,201 +0,0 @@
1
- ---
2
- name: scvi-tools
3
- description: Deep generative models for single-cell omics. Use when you need probabilistic batch correction (scVI), transfer learning, differential expression with uncertainty, or multi-modal integration (TOTALVI, MultiVI). Best for advanced modeling, batch effects, multimodal data. For standard analysis pipelines use scanpy.
4
- license: BSD-3-Clause license
5
- metadata:
6
- version: "1.1"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # scvi-tools
11
-
12
- ## Overview
13
-
14
- scvi-tools is a comprehensive Python framework for probabilistic models in single-cell genomics. Built on PyTorch and PyTorch Lightning, it provides deep generative models using variational inference for analyzing diverse single-cell data modalities. Current stable release: **scvi-tools 1.4.3** (May 2026).
15
-
16
- **Model namespaces matter:** core models (scVI, scANVI, totalVI, MultiVI, PeakVI, AUTOZI, CondSCVI, DestVI, LinearSCVI, AmortizedLDA, JaxSCVI) live under `scvi.model`. Most other models (VeloVI, contrastiveVI, CellAssign, PoissonVI, scBasset, MrVI, MethylVI/MethylANVI, CytoVI, SysVI, Decipher, gimVI, scVIVA, ResolVI, Stereoscope, Solo, totalANVI, DIAGVI) live under `scvi.external`. The reference files specify the correct namespace per model.
17
-
18
- ## When to Use This Skill
19
-
20
- Use this skill when:
21
- - Analyzing single-cell RNA-seq data (dimensionality reduction, batch correction, integration)
22
- - Working with single-cell ATAC-seq or chromatin accessibility data
23
- - Integrating multimodal data (CITE-seq, multiome, paired/unpaired datasets)
24
- - Analyzing spatial transcriptomics data (deconvolution, spatial mapping)
25
- - Performing differential expression analysis on single-cell data
26
- - Conducting cell type annotation or transfer learning tasks
27
- - Working with specialized single-cell modalities (methylation, cytometry, RNA velocity)
28
- - Building custom probabilistic models for single-cell analysis
29
-
30
- ## Core Capabilities
31
-
32
- scvi-tools provides models organized by data modality:
33
-
34
- ### 1. Single-Cell RNA-seq Analysis
35
- Core models for expression analysis, batch correction, and integration. See `references/models-scrna-seq.md` for:
36
- - **scVI**: Unsupervised dimensionality reduction and batch correction
37
- - **scANVI**: Semi-supervised cell type annotation and integration
38
- - **AUTOZI**: Zero-inflation detection and modeling
39
- - **VeloVI**: RNA velocity analysis
40
- - **contrastiveVI**: Perturbation effect isolation
41
-
42
- ### 2. Chromatin Accessibility (ATAC-seq)
43
- Models for analyzing single-cell chromatin data. See `references/models-atac-seq.md` for:
44
- - **PeakVI**: Peak-based ATAC-seq analysis and integration
45
- - **PoissonVI**: Quantitative fragment count modeling
46
- - **scBasset**: Deep learning approach with motif analysis
47
-
48
- ### 3. Multimodal & Multi-omics Integration
49
- Joint analysis of multiple data types. See `references/models-multimodal.md` for:
50
- - **totalVI**: CITE-seq protein and RNA joint modeling
51
- - **totalANVI**: Semi-supervised CITE-seq (totalVI with cell-type labels)
52
- - **MultiVI**: Paired and unpaired multi-omic integration (MuData-based)
53
- - **MrVI**: Multi-resolution cross-sample analysis
54
- - **DIAGVI**: Diagonal integration of unpaired single-cell datasets (added in 1.4.3)
55
-
56
- ### 4. Spatial Transcriptomics
57
- Spatially-resolved transcriptomics analysis. See `references/models-spatial.md` for:
58
- - **DestVI**: Multi-resolution spatial deconvolution
59
- - **Stereoscope**: Cell type deconvolution
60
- - **Tangram**: Spatial mapping and integration
61
- - **scVIVA**: Cell-environment relationship analysis
62
-
63
- ### 5. Specialized Modalities
64
- Additional specialized analysis tools. See `references/models-specialized.md` for:
65
- - **MethylVI/MethylANVI**: Single-cell methylation analysis
66
- - **CytoVI**: Flow/mass cytometry batch correction
67
- - **Solo**: Doublet detection
68
- - **CellAssign**: Marker-based cell type annotation
69
-
70
- ## Typical Workflow
71
-
72
- All scvi-tools models follow a consistent API pattern:
73
-
74
- ```python
75
- # 1. Load and preprocess data (AnnData format)
76
- import scvi
77
- import scanpy as sc
78
-
79
- adata = scvi.data.heart_cell_atlas_subsampled()
80
- sc.pp.filter_genes(adata, min_counts=3)
81
- sc.pp.highly_variable_genes(adata, n_top_genes=1200)
82
-
83
- # 2. Register data with model (specify layers, covariates)
84
- scvi.model.SCVI.setup_anndata(
85
- adata,
86
- layer="counts", # Use raw counts, not log-normalized
87
- batch_key="batch",
88
- categorical_covariate_keys=["donor"],
89
- continuous_covariate_keys=["percent_mito"]
90
- )
91
-
92
- # 3. Create and train model
93
- model = scvi.model.SCVI(adata)
94
- model.train()
95
-
96
- # 4. Extract latent representations and normalized values
97
- latent = model.get_latent_representation()
98
- normalized = model.get_normalized_expression(library_size=1e4)
99
-
100
- # 5. Store in AnnData for downstream analysis
101
- adata.obsm["X_scVI"] = latent
102
- adata.layers["scvi_normalized"] = normalized
103
-
104
- # 6. Downstream analysis with scanpy
105
- sc.pp.neighbors(adata, use_rep="X_scVI")
106
- sc.tl.umap(adata)
107
- sc.tl.leiden(adata)
108
- ```
109
-
110
- **Key Design Principles:**
111
- - **Raw counts required**: Models expect unnormalized count data for optimal performance
112
- - **Unified API**: Consistent interface across all models (setup → train → extract)
113
- - **AnnData-centric**: Seamless integration with the scanpy ecosystem
114
- - **GPU acceleration**: Automatic utilization of available GPUs
115
- - **Batch correction**: Handle technical variation through covariate registration
116
-
117
- ## Common Analysis Tasks
118
-
119
- ### Differential Expression
120
- Probabilistic DE analysis using the learned generative models:
121
-
122
- ```python
123
- de_results = model.differential_expression(
124
- groupby="cell_type",
125
- group1="TypeA",
126
- group2="TypeB",
127
- mode="change", # Use composite hypothesis testing
128
- delta=0.25 # Minimum effect size threshold
129
- )
130
- ```
131
-
132
- See `references/differential-expression.md` for detailed methodology and interpretation.
133
-
134
- ### Model Persistence
135
- Save and load trained models:
136
-
137
- ```python
138
- # Save model
139
- model.save("./model_directory", overwrite=True)
140
-
141
- # Load model
142
- model = scvi.model.SCVI.load("./model_directory", adata=adata)
143
- ```
144
-
145
- ### Batch Correction and Integration
146
- Integrate datasets across batches or studies:
147
-
148
- ```python
149
- # Register batch information
150
- scvi.model.SCVI.setup_anndata(adata, batch_key="study")
151
-
152
- # Model automatically learns batch-corrected representations
153
- model = scvi.model.SCVI(adata)
154
- model.train()
155
- latent = model.get_latent_representation() # Batch-corrected
156
- ```
157
-
158
- ## Theoretical Foundations
159
-
160
- scvi-tools is built on:
161
- - **Variational inference**: Approximate posterior distributions for scalable Bayesian inference
162
- - **Deep generative models**: VAE architectures that learn complex data distributions
163
- - **Amortized inference**: Shared neural networks for efficient learning across cells
164
- - **Probabilistic modeling**: Principled uncertainty quantification and statistical testing
165
-
166
- See `references/theoretical-foundations.md` for detailed background on the mathematical framework.
167
-
168
- ## Additional Resources
169
-
170
- - **Workflows**: `references/workflows.md` contains common workflows, best practices, hyperparameter tuning, and GPU optimization
171
- - **Model References**: Detailed documentation for each model category in the `references/` directory
172
- - **Official Documentation**: https://docs.scvi-tools.org/en/stable/
173
- - **Tutorials**: https://docs.scvi-tools.org/en/stable/tutorials/index.html
174
- - **API Reference**: https://docs.scvi-tools.org/en/stable/api/index.html
175
-
176
- ## Installation
177
-
178
- Requires Python **3.12+** (scvi-tools 1.4 dropped older versions).
179
-
180
- ```bash
181
- uv pip install scvi-tools
182
- # For GPU support
183
- uv pip install "scvi-tools[cuda]"
184
- ```
185
-
186
- For reproducible environments, pin a version: `uv pip install scvi-tools==1.4.3`.
187
-
188
- **Compute backends:** training defaults to PyTorch (CPU/GPU/TPU). A JAX backend
189
- (`scvi.model.JaxSCVI`) and an experimental MLX backend for Apple silicon
190
- (`scvi.model.mlxSCVI`) are available for select models.
191
-
192
- ## Best Practices
193
-
194
- 1. **Use raw counts**: Always provide unnormalized count data to models
195
- 2. **Filter genes**: Remove low-count genes before analysis (e.g., `min_counts=3`)
196
- 3. **Register covariates**: Include known technical factors (batch, donor, etc.) in `setup_anndata`
197
- 4. **Feature selection**: Use highly variable genes for improved performance
198
- 5. **Model saving**: Always save trained models to avoid retraining
199
- 6. **GPU usage**: Enable GPU acceleration for large datasets (`accelerator="gpu"`)
200
- 7. **Scanpy integration**: Store outputs in AnnData objects for downstream analysis
201
-