@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
package/skills/scvelo/SKILL.md
DELETED
|
@@ -1,328 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: scvelo
|
|
3
|
-
description: RNA velocity analysis with scVelo. Estimate cell state transitions from unspliced/spliced mRNA dynamics, infer trajectory directions, compute latent time, and identify driver genes in single-cell RNA-seq data. Complements Scanpy/scVI-tools for trajectory inference.
|
|
4
|
-
license: BSD-3-Clause
|
|
5
|
-
compatibility: Requires Python 3.10+ with scvelo, scanpy, and anndata. Verified against scvelo 0.3.4, whose dynamical model and pl.scatter need pandas<3 and whose stochastic estimator needs numpy<2; the deterministic estimator works on current releases.
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.2"
|
|
8
|
-
skill-author: Kuan-lin Huang
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# scVelo — RNA Velocity Analysis
|
|
12
|
-
|
|
13
|
-
## Overview
|
|
14
|
-
|
|
15
|
-
scVelo is the leading Python package for RNA velocity analysis in single-cell RNA-seq data. It infers cell state transitions by modeling the kinetics of mRNA splicing — using the ratio of unspliced (pre-mRNA) to spliced (mature mRNA) abundances to determine whether a gene is being upregulated or downregulated in each cell. This allows reconstruction of developmental trajectories and identification of cell fate decisions without requiring time-course data.
|
|
16
|
-
|
|
17
|
-
**Installation:** `uv pip install scvelo`
|
|
18
|
-
|
|
19
|
-
**Key resources:**
|
|
20
|
-
- Documentation: https://scvelo.readthedocs.io/
|
|
21
|
-
- GitHub: https://github.com/theislab/scvelo
|
|
22
|
-
- Paper: Bergen et al. (2020) Nature Biotechnology. PMID: 32747759
|
|
23
|
-
|
|
24
|
-
## When to Use This Skill
|
|
25
|
-
|
|
26
|
-
Use scVelo when:
|
|
27
|
-
|
|
28
|
-
- **Trajectory inference from snapshot data**: Determine which direction cells are differentiating
|
|
29
|
-
- **Cell fate prediction**: Identify progenitor cells and their downstream fates
|
|
30
|
-
- **Driver gene identification**: Find genes whose dynamics best explain observed trajectories
|
|
31
|
-
- **Developmental biology**: Model hematopoiesis, neurogenesis, epithelial-to-mesenchymal transitions
|
|
32
|
-
- **Latent time estimation**: Order cells along a pseudotime derived from splicing dynamics
|
|
33
|
-
- **Complement to Scanpy**: Add directional information to UMAP embeddings
|
|
34
|
-
|
|
35
|
-
## Prerequisites
|
|
36
|
-
|
|
37
|
-
scVelo requires count matrices for both **unspliced** and **spliced** RNA. These are generated by:
|
|
38
|
-
1. **STARsolo** or **kallisto|bustools** with `lamanno` mode
|
|
39
|
-
2. **velocyto** CLI: `velocyto run10x` / `velocyto run`
|
|
40
|
-
3. **alevin-fry** / **simpleaf** with spliced/unspliced output
|
|
41
|
-
|
|
42
|
-
Data is stored in an `AnnData` object with `layers["spliced"]` and `layers["unspliced"]`.
|
|
43
|
-
|
|
44
|
-
## Standard RNA Velocity Workflow
|
|
45
|
-
|
|
46
|
-
### 1. Setup and Data Loading
|
|
47
|
-
|
|
48
|
-
```python
|
|
49
|
-
import scvelo as scv
|
|
50
|
-
import scanpy as sc
|
|
51
|
-
import numpy as np
|
|
52
|
-
import matplotlib.pyplot as plt
|
|
53
|
-
|
|
54
|
-
# Configure settings
|
|
55
|
-
scv.settings.verbosity = 3 # Show computation steps
|
|
56
|
-
scv.settings.presenter_view = True
|
|
57
|
-
scv.settings.set_figure_params('scvelo')
|
|
58
|
-
|
|
59
|
-
# Load data (AnnData with spliced/unspliced layers)
|
|
60
|
-
# Option A: Load from loom (velocyto output)
|
|
61
|
-
adata = scv.read("cellranger_output.loom", cache=True)
|
|
62
|
-
|
|
63
|
-
# Option B: Merge velocyto loom with Scanpy-processed AnnData
|
|
64
|
-
adata_processed = sc.read_h5ad("processed.h5ad") # Has UMAP, clusters
|
|
65
|
-
adata_velocity = scv.read("velocyto.loom")
|
|
66
|
-
adata = scv.utils.merge(adata_processed, adata_velocity)
|
|
67
|
-
|
|
68
|
-
# Verify layers
|
|
69
|
-
print(adata)
|
|
70
|
-
# obs × var: N × G
|
|
71
|
-
# layers: 'spliced', 'unspliced' (required)
|
|
72
|
-
# obsm['X_umap'] (required for visualization)
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
### 2. Preprocessing
|
|
76
|
-
|
|
77
|
-
```python
|
|
78
|
-
# Filter and normalize. As of scVelo 0.3, filter_and_normalize() only filters
|
|
79
|
-
# genes and normalizes per cell -- it no longer takes n_top_genes and no longer
|
|
80
|
-
# log-transforms, so the log step and HVG selection come from Scanpy.
|
|
81
|
-
scv.pp.filter_and_normalize(
|
|
82
|
-
adata,
|
|
83
|
-
min_shared_counts=20 # Minimum counts in spliced+unspliced
|
|
84
|
-
)
|
|
85
|
-
sc.pp.log1p(adata)
|
|
86
|
-
sc.pp.highly_variable_genes(adata, n_top_genes=2000, subset=True)
|
|
87
|
-
|
|
88
|
-
# Compute first and second order moments (means and variances)
|
|
89
|
-
# knn_connectivities must be computed first
|
|
90
|
-
sc.pp.neighbors(adata, n_neighbors=30, n_pcs=30)
|
|
91
|
-
scv.pp.moments(
|
|
92
|
-
adata,
|
|
93
|
-
n_pcs=30,
|
|
94
|
-
n_neighbors=30
|
|
95
|
-
)
|
|
96
|
-
```
|
|
97
|
-
|
|
98
|
-
### 3. Velocity Estimation — Stochastic Model
|
|
99
|
-
|
|
100
|
-
The stochastic model is fast and suitable for exploratory analysis:
|
|
101
|
-
|
|
102
|
-
```python
|
|
103
|
-
# Stochastic velocity (faster, less accurate)
|
|
104
|
-
scv.tl.velocity(adata, mode='stochastic')
|
|
105
|
-
scv.tl.velocity_graph(adata)
|
|
106
|
-
|
|
107
|
-
# Visualize
|
|
108
|
-
scv.pl.velocity_embedding_stream(
|
|
109
|
-
adata,
|
|
110
|
-
basis='umap',
|
|
111
|
-
color='leiden',
|
|
112
|
-
title="RNA Velocity (Stochastic)"
|
|
113
|
-
)
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
### 4. Velocity Estimation — Dynamical Model (Recommended)
|
|
117
|
-
|
|
118
|
-
The dynamical model fits the full splicing kinetics and is more accurate:
|
|
119
|
-
|
|
120
|
-
```python
|
|
121
|
-
# Recover dynamics (computationally intensive; ~10-30 min for 10K cells)
|
|
122
|
-
scv.tl.recover_dynamics(adata, n_jobs=4)
|
|
123
|
-
|
|
124
|
-
# Compute velocity from dynamical model
|
|
125
|
-
scv.tl.velocity(adata, mode='dynamical')
|
|
126
|
-
scv.tl.velocity_graph(adata)
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
### 5. Latent Time
|
|
130
|
-
|
|
131
|
-
The dynamical model enables computation of a shared latent time (pseudotime):
|
|
132
|
-
|
|
133
|
-
```python
|
|
134
|
-
# Compute latent time
|
|
135
|
-
scv.tl.latent_time(adata)
|
|
136
|
-
|
|
137
|
-
# Visualize latent time on UMAP
|
|
138
|
-
scv.pl.scatter(
|
|
139
|
-
adata,
|
|
140
|
-
color='latent_time',
|
|
141
|
-
color_map='gnuplot',
|
|
142
|
-
size=80,
|
|
143
|
-
title='Latent time'
|
|
144
|
-
)
|
|
145
|
-
|
|
146
|
-
# Identify top genes ordered by latent time
|
|
147
|
-
top_genes = adata.var['fit_likelihood'].sort_values(ascending=False).index[:300]
|
|
148
|
-
scv.pl.heatmap(
|
|
149
|
-
adata,
|
|
150
|
-
var_names=top_genes,
|
|
151
|
-
sortby='latent_time',
|
|
152
|
-
col_color='leiden',
|
|
153
|
-
n_convolve=100
|
|
154
|
-
)
|
|
155
|
-
```
|
|
156
|
-
|
|
157
|
-
### 6. Driver Gene Analysis
|
|
158
|
-
|
|
159
|
-
```python
|
|
160
|
-
# Identify genes with highest velocity fit
|
|
161
|
-
scv.tl.rank_velocity_genes(adata, groupby='leiden', min_corr=0.3)
|
|
162
|
-
df = scv.DataFrame(adata.uns['rank_velocity_genes']['names'])
|
|
163
|
-
print(df.head(10))
|
|
164
|
-
|
|
165
|
-
# Speed and coherence
|
|
166
|
-
scv.tl.velocity_confidence(adata)
|
|
167
|
-
scv.pl.scatter(
|
|
168
|
-
adata,
|
|
169
|
-
c=['velocity_length', 'velocity_confidence'],
|
|
170
|
-
cmap='coolwarm',
|
|
171
|
-
perc=[5, 95]
|
|
172
|
-
)
|
|
173
|
-
|
|
174
|
-
# Phase portraits for specific genes
|
|
175
|
-
scv.pl.velocity(adata, ['Cpe', 'Gnao1', 'Ins2'],
|
|
176
|
-
ncols=3, figsize=(16, 4))
|
|
177
|
-
```
|
|
178
|
-
|
|
179
|
-
### 7. Velocity Arrows and Pseudotime
|
|
180
|
-
|
|
181
|
-
```python
|
|
182
|
-
# Arrow plot on UMAP
|
|
183
|
-
scv.pl.velocity_embedding(
|
|
184
|
-
adata,
|
|
185
|
-
arrow_length=3,
|
|
186
|
-
arrow_size=2,
|
|
187
|
-
color='leiden',
|
|
188
|
-
basis='umap'
|
|
189
|
-
)
|
|
190
|
-
|
|
191
|
-
# Stream plot (cleaner visualization)
|
|
192
|
-
scv.pl.velocity_embedding_stream(
|
|
193
|
-
adata,
|
|
194
|
-
basis='umap',
|
|
195
|
-
color='leiden',
|
|
196
|
-
smooth=0.8,
|
|
197
|
-
min_mass=4
|
|
198
|
-
)
|
|
199
|
-
|
|
200
|
-
# Velocity pseudotime (alternative to latent time)
|
|
201
|
-
scv.tl.velocity_pseudotime(adata)
|
|
202
|
-
scv.pl.scatter(adata, color='velocity_pseudotime', cmap='gnuplot')
|
|
203
|
-
```
|
|
204
|
-
|
|
205
|
-
### 8. PAGA Trajectory Graph
|
|
206
|
-
|
|
207
|
-
```python
|
|
208
|
-
# PAGA graph with velocity-informed transitions
|
|
209
|
-
scv.tl.paga(adata, groups='leiden')
|
|
210
|
-
df = scv.get_df(adata, 'paga/transitions_confidence', precision=2).T
|
|
211
|
-
df.style.background_gradient(cmap='Blues').format('{:.2g}')
|
|
212
|
-
|
|
213
|
-
# Plot PAGA with velocity
|
|
214
|
-
scv.pl.paga(
|
|
215
|
-
adata,
|
|
216
|
-
basis='umap',
|
|
217
|
-
size=50,
|
|
218
|
-
alpha=0.1,
|
|
219
|
-
min_edge_width=2,
|
|
220
|
-
node_size_scale=1.5
|
|
221
|
-
)
|
|
222
|
-
```
|
|
223
|
-
|
|
224
|
-
## Complete Workflow Script
|
|
225
|
-
|
|
226
|
-
```python
|
|
227
|
-
import scvelo as scv
|
|
228
|
-
import scanpy as sc
|
|
229
|
-
|
|
230
|
-
def run_rna_velocity(adata, n_top_genes=2000, mode='dynamical', n_jobs=4):
|
|
231
|
-
"""
|
|
232
|
-
Complete RNA velocity workflow.
|
|
233
|
-
|
|
234
|
-
Args:
|
|
235
|
-
adata: AnnData with 'spliced' and 'unspliced' layers, UMAP in obsm
|
|
236
|
-
n_top_genes: Number of top HVGs for velocity
|
|
237
|
-
mode: 'stochastic' (fast) or 'dynamical' (accurate)
|
|
238
|
-
n_jobs: Parallel jobs for dynamical model
|
|
239
|
-
|
|
240
|
-
Returns:
|
|
241
|
-
Processed AnnData with velocity information
|
|
242
|
-
"""
|
|
243
|
-
scv.settings.verbosity = 2
|
|
244
|
-
|
|
245
|
-
# 1. Preprocessing (scVelo 0.3 dropped log/HVG from filter_and_normalize)
|
|
246
|
-
scv.pp.filter_and_normalize(adata, min_shared_counts=20)
|
|
247
|
-
sc.pp.log1p(adata)
|
|
248
|
-
sc.pp.highly_variable_genes(adata, n_top_genes=n_top_genes, subset=True)
|
|
249
|
-
|
|
250
|
-
if 'neighbors' not in adata.uns:
|
|
251
|
-
sc.pp.neighbors(adata, n_neighbors=30)
|
|
252
|
-
|
|
253
|
-
scv.pp.moments(adata, n_pcs=30, n_neighbors=30)
|
|
254
|
-
|
|
255
|
-
# 2. Velocity estimation
|
|
256
|
-
if mode == 'dynamical':
|
|
257
|
-
scv.tl.recover_dynamics(adata, n_jobs=n_jobs)
|
|
258
|
-
|
|
259
|
-
scv.tl.velocity(adata, mode=mode)
|
|
260
|
-
scv.tl.velocity_graph(adata)
|
|
261
|
-
|
|
262
|
-
# 3. Downstream analyses
|
|
263
|
-
if mode == 'dynamical':
|
|
264
|
-
scv.tl.latent_time(adata)
|
|
265
|
-
scv.tl.rank_velocity_genes(adata, groupby='leiden', min_corr=0.3)
|
|
266
|
-
|
|
267
|
-
scv.tl.velocity_confidence(adata)
|
|
268
|
-
scv.tl.velocity_pseudotime(adata)
|
|
269
|
-
|
|
270
|
-
return adata
|
|
271
|
-
```
|
|
272
|
-
|
|
273
|
-
## Key Output Fields in AnnData
|
|
274
|
-
|
|
275
|
-
After running the workflow, the following fields are added:
|
|
276
|
-
|
|
277
|
-
| Location | Key | Description |
|
|
278
|
-
|----------|-----|-------------|
|
|
279
|
-
| `adata.layers` | `velocity` | RNA velocity per gene per cell |
|
|
280
|
-
| `adata.layers` | `fit_t` | Fitted latent time per gene per cell |
|
|
281
|
-
| `adata.obsm` | `velocity_umap` | 2D velocity vectors on UMAP |
|
|
282
|
-
| `adata.obs` | `velocity_pseudotime` | Pseudotime from velocity |
|
|
283
|
-
| `adata.obs` | `latent_time` | Latent time from dynamical model |
|
|
284
|
-
| `adata.obs` | `velocity_length` | Speed of each cell |
|
|
285
|
-
| `adata.obs` | `velocity_confidence` | Confidence score per cell |
|
|
286
|
-
| `adata.var` | `fit_likelihood` | Gene-level model fit quality |
|
|
287
|
-
| `adata.var` | `fit_alpha` | Transcription rate |
|
|
288
|
-
| `adata.var` | `fit_beta` | Splicing rate |
|
|
289
|
-
| `adata.var` | `fit_gamma` | Degradation rate |
|
|
290
|
-
| `adata.uns` | `velocity_graph` | Cell-cell transition probability matrix |
|
|
291
|
-
|
|
292
|
-
## Velocity Models Comparison
|
|
293
|
-
|
|
294
|
-
| Model | Speed | Accuracy | When to Use |
|
|
295
|
-
|-------|-------|----------|-------------|
|
|
296
|
-
| `stochastic` | Fast | Moderate | Exploratory; large datasets |
|
|
297
|
-
| `deterministic` | Medium | Moderate | Simple linear kinetics |
|
|
298
|
-
| `dynamical` | Slow | High | Publication-quality; identifies driver genes |
|
|
299
|
-
|
|
300
|
-
## Best Practices
|
|
301
|
-
|
|
302
|
-
- **Start with stochastic mode** for exploration; switch to dynamical for final analysis
|
|
303
|
-
- **Need good coverage of unspliced reads**: Short reads (< 100 bp) may miss intron coverage
|
|
304
|
-
- **Minimum 2,000 cells**: RNA velocity is noisy with fewer cells
|
|
305
|
-
- **Velocity should be coherent**: Arrows should follow known biology; randomness indicates issues
|
|
306
|
-
- **k-NN bandwidth matters**: Too few neighbors → noisy velocity; too many → oversmoothed
|
|
307
|
-
- **Sanity check**: Root cells (progenitors) should have high unspliced/spliced ratios for marker genes
|
|
308
|
-
- **Dynamical model requires distinct kinetic states**: Works best for clear differentiation processes
|
|
309
|
-
|
|
310
|
-
## Troubleshooting
|
|
311
|
-
|
|
312
|
-
| Problem | Solution |
|
|
313
|
-
|---------|---------|
|
|
314
|
-
| Missing unspliced layer | Re-run velocyto or use STARsolo with `--soloFeatures Gene Velocyto` |
|
|
315
|
-
| Very few velocity genes | Lower `min_shared_counts`; check sequencing depth |
|
|
316
|
-
| Random-looking arrows | Try different `n_neighbors` or velocity model |
|
|
317
|
-
| Memory error with dynamical | Set `n_jobs=1`; reduce `n_top_genes` |
|
|
318
|
-
| Negative velocity everywhere | Check that spliced/unspliced layers are not swapped |
|
|
319
|
-
|
|
320
|
-
## Additional Resources
|
|
321
|
-
|
|
322
|
-
- **scVelo documentation**: https://scvelo.readthedocs.io/
|
|
323
|
-
- **Tutorial notebooks**: https://scvelo.readthedocs.io/tutorials/
|
|
324
|
-
- **GitHub**: https://github.com/theislab/scvelo
|
|
325
|
-
- **Paper**: Bergen V et al. (2020) Nature Biotechnology. PMID: 32747759
|
|
326
|
-
- **velocyto** (preprocessing): http://velocyto.org/
|
|
327
|
-
- **CellRank** (fate prediction, extends scVelo): https://cellrank.readthedocs.io/
|
|
328
|
-
- **dynamo** (metabolic labeling alternative): https://dynamo-release.readthedocs.io/
|
|
@@ -1,201 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: scvi-tools
|
|
3
|
-
description: Deep generative models for single-cell omics. Use when you need probabilistic batch correction (scVI), transfer learning, differential expression with uncertainty, or multi-modal integration (TOTALVI, MultiVI). Best for advanced modeling, batch effects, multimodal data. For standard analysis pipelines use scanpy.
|
|
4
|
-
license: BSD-3-Clause license
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.1"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# scvi-tools
|
|
11
|
-
|
|
12
|
-
## Overview
|
|
13
|
-
|
|
14
|
-
scvi-tools is a comprehensive Python framework for probabilistic models in single-cell genomics. Built on PyTorch and PyTorch Lightning, it provides deep generative models using variational inference for analyzing diverse single-cell data modalities. Current stable release: **scvi-tools 1.4.3** (May 2026).
|
|
15
|
-
|
|
16
|
-
**Model namespaces matter:** core models (scVI, scANVI, totalVI, MultiVI, PeakVI, AUTOZI, CondSCVI, DestVI, LinearSCVI, AmortizedLDA, JaxSCVI) live under `scvi.model`. Most other models (VeloVI, contrastiveVI, CellAssign, PoissonVI, scBasset, MrVI, MethylVI/MethylANVI, CytoVI, SysVI, Decipher, gimVI, scVIVA, ResolVI, Stereoscope, Solo, totalANVI, DIAGVI) live under `scvi.external`. The reference files specify the correct namespace per model.
|
|
17
|
-
|
|
18
|
-
## When to Use This Skill
|
|
19
|
-
|
|
20
|
-
Use this skill when:
|
|
21
|
-
- Analyzing single-cell RNA-seq data (dimensionality reduction, batch correction, integration)
|
|
22
|
-
- Working with single-cell ATAC-seq or chromatin accessibility data
|
|
23
|
-
- Integrating multimodal data (CITE-seq, multiome, paired/unpaired datasets)
|
|
24
|
-
- Analyzing spatial transcriptomics data (deconvolution, spatial mapping)
|
|
25
|
-
- Performing differential expression analysis on single-cell data
|
|
26
|
-
- Conducting cell type annotation or transfer learning tasks
|
|
27
|
-
- Working with specialized single-cell modalities (methylation, cytometry, RNA velocity)
|
|
28
|
-
- Building custom probabilistic models for single-cell analysis
|
|
29
|
-
|
|
30
|
-
## Core Capabilities
|
|
31
|
-
|
|
32
|
-
scvi-tools provides models organized by data modality:
|
|
33
|
-
|
|
34
|
-
### 1. Single-Cell RNA-seq Analysis
|
|
35
|
-
Core models for expression analysis, batch correction, and integration. See `references/models-scrna-seq.md` for:
|
|
36
|
-
- **scVI**: Unsupervised dimensionality reduction and batch correction
|
|
37
|
-
- **scANVI**: Semi-supervised cell type annotation and integration
|
|
38
|
-
- **AUTOZI**: Zero-inflation detection and modeling
|
|
39
|
-
- **VeloVI**: RNA velocity analysis
|
|
40
|
-
- **contrastiveVI**: Perturbation effect isolation
|
|
41
|
-
|
|
42
|
-
### 2. Chromatin Accessibility (ATAC-seq)
|
|
43
|
-
Models for analyzing single-cell chromatin data. See `references/models-atac-seq.md` for:
|
|
44
|
-
- **PeakVI**: Peak-based ATAC-seq analysis and integration
|
|
45
|
-
- **PoissonVI**: Quantitative fragment count modeling
|
|
46
|
-
- **scBasset**: Deep learning approach with motif analysis
|
|
47
|
-
|
|
48
|
-
### 3. Multimodal & Multi-omics Integration
|
|
49
|
-
Joint analysis of multiple data types. See `references/models-multimodal.md` for:
|
|
50
|
-
- **totalVI**: CITE-seq protein and RNA joint modeling
|
|
51
|
-
- **totalANVI**: Semi-supervised CITE-seq (totalVI with cell-type labels)
|
|
52
|
-
- **MultiVI**: Paired and unpaired multi-omic integration (MuData-based)
|
|
53
|
-
- **MrVI**: Multi-resolution cross-sample analysis
|
|
54
|
-
- **DIAGVI**: Diagonal integration of unpaired single-cell datasets (added in 1.4.3)
|
|
55
|
-
|
|
56
|
-
### 4. Spatial Transcriptomics
|
|
57
|
-
Spatially-resolved transcriptomics analysis. See `references/models-spatial.md` for:
|
|
58
|
-
- **DestVI**: Multi-resolution spatial deconvolution
|
|
59
|
-
- **Stereoscope**: Cell type deconvolution
|
|
60
|
-
- **Tangram**: Spatial mapping and integration
|
|
61
|
-
- **scVIVA**: Cell-environment relationship analysis
|
|
62
|
-
|
|
63
|
-
### 5. Specialized Modalities
|
|
64
|
-
Additional specialized analysis tools. See `references/models-specialized.md` for:
|
|
65
|
-
- **MethylVI/MethylANVI**: Single-cell methylation analysis
|
|
66
|
-
- **CytoVI**: Flow/mass cytometry batch correction
|
|
67
|
-
- **Solo**: Doublet detection
|
|
68
|
-
- **CellAssign**: Marker-based cell type annotation
|
|
69
|
-
|
|
70
|
-
## Typical Workflow
|
|
71
|
-
|
|
72
|
-
All scvi-tools models follow a consistent API pattern:
|
|
73
|
-
|
|
74
|
-
```python
|
|
75
|
-
# 1. Load and preprocess data (AnnData format)
|
|
76
|
-
import scvi
|
|
77
|
-
import scanpy as sc
|
|
78
|
-
|
|
79
|
-
adata = scvi.data.heart_cell_atlas_subsampled()
|
|
80
|
-
sc.pp.filter_genes(adata, min_counts=3)
|
|
81
|
-
sc.pp.highly_variable_genes(adata, n_top_genes=1200)
|
|
82
|
-
|
|
83
|
-
# 2. Register data with model (specify layers, covariates)
|
|
84
|
-
scvi.model.SCVI.setup_anndata(
|
|
85
|
-
adata,
|
|
86
|
-
layer="counts", # Use raw counts, not log-normalized
|
|
87
|
-
batch_key="batch",
|
|
88
|
-
categorical_covariate_keys=["donor"],
|
|
89
|
-
continuous_covariate_keys=["percent_mito"]
|
|
90
|
-
)
|
|
91
|
-
|
|
92
|
-
# 3. Create and train model
|
|
93
|
-
model = scvi.model.SCVI(adata)
|
|
94
|
-
model.train()
|
|
95
|
-
|
|
96
|
-
# 4. Extract latent representations and normalized values
|
|
97
|
-
latent = model.get_latent_representation()
|
|
98
|
-
normalized = model.get_normalized_expression(library_size=1e4)
|
|
99
|
-
|
|
100
|
-
# 5. Store in AnnData for downstream analysis
|
|
101
|
-
adata.obsm["X_scVI"] = latent
|
|
102
|
-
adata.layers["scvi_normalized"] = normalized
|
|
103
|
-
|
|
104
|
-
# 6. Downstream analysis with scanpy
|
|
105
|
-
sc.pp.neighbors(adata, use_rep="X_scVI")
|
|
106
|
-
sc.tl.umap(adata)
|
|
107
|
-
sc.tl.leiden(adata)
|
|
108
|
-
```
|
|
109
|
-
|
|
110
|
-
**Key Design Principles:**
|
|
111
|
-
- **Raw counts required**: Models expect unnormalized count data for optimal performance
|
|
112
|
-
- **Unified API**: Consistent interface across all models (setup → train → extract)
|
|
113
|
-
- **AnnData-centric**: Seamless integration with the scanpy ecosystem
|
|
114
|
-
- **GPU acceleration**: Automatic utilization of available GPUs
|
|
115
|
-
- **Batch correction**: Handle technical variation through covariate registration
|
|
116
|
-
|
|
117
|
-
## Common Analysis Tasks
|
|
118
|
-
|
|
119
|
-
### Differential Expression
|
|
120
|
-
Probabilistic DE analysis using the learned generative models:
|
|
121
|
-
|
|
122
|
-
```python
|
|
123
|
-
de_results = model.differential_expression(
|
|
124
|
-
groupby="cell_type",
|
|
125
|
-
group1="TypeA",
|
|
126
|
-
group2="TypeB",
|
|
127
|
-
mode="change", # Use composite hypothesis testing
|
|
128
|
-
delta=0.25 # Minimum effect size threshold
|
|
129
|
-
)
|
|
130
|
-
```
|
|
131
|
-
|
|
132
|
-
See `references/differential-expression.md` for detailed methodology and interpretation.
|
|
133
|
-
|
|
134
|
-
### Model Persistence
|
|
135
|
-
Save and load trained models:
|
|
136
|
-
|
|
137
|
-
```python
|
|
138
|
-
# Save model
|
|
139
|
-
model.save("./model_directory", overwrite=True)
|
|
140
|
-
|
|
141
|
-
# Load model
|
|
142
|
-
model = scvi.model.SCVI.load("./model_directory", adata=adata)
|
|
143
|
-
```
|
|
144
|
-
|
|
145
|
-
### Batch Correction and Integration
|
|
146
|
-
Integrate datasets across batches or studies:
|
|
147
|
-
|
|
148
|
-
```python
|
|
149
|
-
# Register batch information
|
|
150
|
-
scvi.model.SCVI.setup_anndata(adata, batch_key="study")
|
|
151
|
-
|
|
152
|
-
# Model automatically learns batch-corrected representations
|
|
153
|
-
model = scvi.model.SCVI(adata)
|
|
154
|
-
model.train()
|
|
155
|
-
latent = model.get_latent_representation() # Batch-corrected
|
|
156
|
-
```
|
|
157
|
-
|
|
158
|
-
## Theoretical Foundations
|
|
159
|
-
|
|
160
|
-
scvi-tools is built on:
|
|
161
|
-
- **Variational inference**: Approximate posterior distributions for scalable Bayesian inference
|
|
162
|
-
- **Deep generative models**: VAE architectures that learn complex data distributions
|
|
163
|
-
- **Amortized inference**: Shared neural networks for efficient learning across cells
|
|
164
|
-
- **Probabilistic modeling**: Principled uncertainty quantification and statistical testing
|
|
165
|
-
|
|
166
|
-
See `references/theoretical-foundations.md` for detailed background on the mathematical framework.
|
|
167
|
-
|
|
168
|
-
## Additional Resources
|
|
169
|
-
|
|
170
|
-
- **Workflows**: `references/workflows.md` contains common workflows, best practices, hyperparameter tuning, and GPU optimization
|
|
171
|
-
- **Model References**: Detailed documentation for each model category in the `references/` directory
|
|
172
|
-
- **Official Documentation**: https://docs.scvi-tools.org/en/stable/
|
|
173
|
-
- **Tutorials**: https://docs.scvi-tools.org/en/stable/tutorials/index.html
|
|
174
|
-
- **API Reference**: https://docs.scvi-tools.org/en/stable/api/index.html
|
|
175
|
-
|
|
176
|
-
## Installation
|
|
177
|
-
|
|
178
|
-
Requires Python **3.12+** (scvi-tools 1.4 dropped older versions).
|
|
179
|
-
|
|
180
|
-
```bash
|
|
181
|
-
uv pip install scvi-tools
|
|
182
|
-
# For GPU support
|
|
183
|
-
uv pip install "scvi-tools[cuda]"
|
|
184
|
-
```
|
|
185
|
-
|
|
186
|
-
For reproducible environments, pin a version: `uv pip install scvi-tools==1.4.3`.
|
|
187
|
-
|
|
188
|
-
**Compute backends:** training defaults to PyTorch (CPU/GPU/TPU). A JAX backend
|
|
189
|
-
(`scvi.model.JaxSCVI`) and an experimental MLX backend for Apple silicon
|
|
190
|
-
(`scvi.model.mlxSCVI`) are available for select models.
|
|
191
|
-
|
|
192
|
-
## Best Practices
|
|
193
|
-
|
|
194
|
-
1. **Use raw counts**: Always provide unnormalized count data to models
|
|
195
|
-
2. **Filter genes**: Remove low-count genes before analysis (e.g., `min_counts=3`)
|
|
196
|
-
3. **Register covariates**: Include known technical factors (batch, donor, etc.) in `setup_anndata`
|
|
197
|
-
4. **Feature selection**: Use highly variable genes for improved performance
|
|
198
|
-
5. **Model saving**: Always save trained models to avoid retraining
|
|
199
|
-
6. **GPU usage**: Enable GPU acceleration for large datasets (`accelerator="gpu"`)
|
|
200
|
-
7. **Scanpy integration**: Store outputs in AnnData objects for downstream analysis
|
|
201
|
-
|