@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,244 +0,0 @@
1
- ---
2
- name: deepchem
3
- description: Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featurization options and MoleculeNet benchmarks. Best for quick experiments with pre-trained models, diverse molecular representations. For graph-first PyTorch workflows use torchdrug; for benchmark datasets use pytdc.
4
- license: MIT license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.7–3.11 (PyPI 2.8.0 caps at <3.12). Install PyTorch, TensorFlow, or JAX before the matching deepchem extra. RDKit is a core dependency.
7
- metadata:
8
- version: "1.4"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # DeepChem
13
-
14
- ## Overview
15
-
16
- DeepChem is a comprehensive Python library for applying machine learning to chemistry, materials science, and biology. Enable molecular property prediction, drug discovery, materials design, and biomolecule analysis through specialized neural networks, molecular featurization methods, and pretrained models.
17
-
18
- **Version note:** Examples target **deepchem 2.8.0** (PyPI stable, Apr 2024). Requires **Python 3.7–3.11** (`<3.12` on PyPI). Core utilities (loaders, featurizers, MoleculeNet) work without a DL backend; GNN and transformer models need the matching extra (`torch`, `tensorflow`, or `jax`). Install the backend framework first when using GPU builds.
19
-
20
- ## When to Use This Skill
21
-
22
- This skill should be used when:
23
- - Loading and processing molecular data (SMILES strings, SDF files, protein sequences)
24
- - Predicting molecular properties (solubility, toxicity, binding affinity, ADMET properties)
25
- - Training models on chemical/biological datasets
26
- - Using MoleculeNet benchmark datasets (Tox21, BBBP, Delaney, etc.)
27
- - Converting molecules to ML-ready features (fingerprints, graph representations, descriptors)
28
- - Implementing graph neural networks for molecules (GCN, GAT, MPNN, AttentiveFP)
29
- - Applying transfer learning with pretrained models (ChemBERTa, GROVER, MolFormer)
30
- - Predicting crystal/materials properties (bandgap, formation energy)
31
- - Analyzing protein or DNA sequences
32
-
33
- ## Core Capabilities
34
-
35
- Eight capability areas, each with worked code, are in
36
- [references/core_capabilities.md](references/core_capabilities.md):
37
-
38
- 1. **Molecular data loading and processing** — loaders, `NumpyDataset` / `DiskDataset`.
39
- 2. **Molecular featurization** — circular fingerprints, graph convolution, and descriptors.
40
- 3. **Data splitting** — random, scaffold, stratified, and butina splitters, and why
41
- scaffold splitting is the honest default for molecules.
42
- 4. **Model selection and training** — the model families and how to fit them.
43
- 5. **MoleculeNet benchmarks** — loading standard datasets and their published splits.
44
- 6. **Transfer learning** — pretraining and fine-tuning.
45
- 7. **Model evaluation** — metrics appropriate to regression and classification tasks.
46
- 8. **Making predictions** — applying a trained model to new molecules.
47
-
48
- Three end-to-end workflows are in
49
- [references/typical_workflows.md](references/typical_workflows.md).
50
-
51
- ## Example Scripts
52
-
53
- This skill includes three production-ready scripts in the `scripts/` directory:
54
-
55
- ### 1. `predict_solubility.py`
56
- Train and evaluate solubility prediction models. Works with Delaney benchmark or custom CSV data.
57
-
58
- ```bash
59
- # Use Delaney benchmark
60
- python scripts/predict_solubility.py
61
-
62
- # Use custom data
63
- python scripts/predict_solubility.py \
64
- --data my_data.csv \
65
- --smiles-col smiles \
66
- --target-col solubility \
67
- --predict "CCO" "c1ccccc1"
68
- ```
69
-
70
- ### 2. `graph_neural_network.py`
71
- Train various graph neural network architectures on molecular data.
72
-
73
- ```bash
74
- # Train GCN on Tox21
75
- python scripts/graph_neural_network.py --model gcn --dataset tox21
76
-
77
- # Train AttentiveFP on custom data
78
- python scripts/graph_neural_network.py \
79
- --model attentivefp \
80
- --data molecules.csv \
81
- --task-type regression \
82
- --targets activity \
83
- --epochs 100
84
- ```
85
-
86
- ### 3. `transfer_learning.py`
87
- Fine-tune pretrained models (ChemBERTa, GROVER, MolFormer) on molecular property prediction tasks.
88
-
89
- ```bash
90
- # Fine-tune ChemBERTa on BBBP
91
- python scripts/transfer_learning.py --model chemberta --dataset bbbp
92
-
93
- # Fine-tune GROVER on custom data
94
- python scripts/transfer_learning.py \
95
- --model grover \
96
- --data small_dataset.csv \
97
- --target activity \
98
- --task-type classification \
99
- --epochs 20
100
- ```
101
-
102
- ## Common Patterns and Best Practices
103
-
104
- ### Pattern 1: Always Use Scaffold Splitting for Molecules
105
- ```python
106
- # GOOD: Prevents data leakage
107
- splitter = dc.splits.ScaffoldSplitter()
108
- train, test = splitter.train_test_split(dataset)
109
-
110
- # BAD: Similar molecules in train and test
111
- splitter = dc.splits.RandomSplitter()
112
- train, test = splitter.train_test_split(dataset)
113
- ```
114
-
115
- ### Pattern 2: Normalize Features and Targets
116
- ```python
117
- transformers = [
118
- dc.trans.NormalizationTransformer(
119
- transform_y=True, # Also normalize target values
120
- dataset=train
121
- )
122
- ]
123
- for transformer in transformers:
124
- train = transformer.transform(train)
125
- test = transformer.transform(test)
126
- ```
127
-
128
- ### Pattern 3: Start Simple, Then Scale
129
- 1. Start with Random Forest + CircularFingerprint (fast baseline)
130
- 2. Try XGBoost/LightGBM if RF works well
131
- 3. Move to deep learning (MultitaskRegressor) if you have >5K samples
132
- 4. Try GNNs if you have >10K samples
133
- 5. Use transfer learning for small datasets or novel scaffolds
134
-
135
- ### Pattern 4: Handle Imbalanced Data
136
- ```python
137
- # Option 1: Balancing transformer
138
- transformer = dc.trans.BalancingTransformer(dataset=train)
139
- train = transformer.transform(train)
140
-
141
- # Option 2: Use balanced metrics
142
- metric = dc.metrics.Metric(dc.metrics.balanced_accuracy_score)
143
- ```
144
-
145
- ### Pattern 5: Avoid Memory Issues
146
- ```python
147
- # Use DiskDataset for large datasets
148
- dataset = dc.data.DiskDataset.from_numpy(X, y, w, ids)
149
-
150
- # Use smaller batch sizes
151
- model = dc.models.GCNModel(batch_size=32) # Instead of 128
152
- ```
153
-
154
- ## Common Pitfalls
155
-
156
- ### Issue 1: Data Leakage in Drug Discovery
157
- **Problem**: Using random splitting allows similar molecules in train/test sets.
158
- **Solution**: Always use `ScaffoldSplitter` for molecular datasets.
159
-
160
- ### Issue 2: GNN Underperforming vs Fingerprints
161
- **Problem**: Graph neural networks perform worse than simple fingerprints.
162
- **Solutions**:
163
- - Ensure dataset is large enough (>10K samples typically)
164
- - Increase training epochs (50-100)
165
- - Try different architectures (AttentiveFP, DMPNN instead of GCN)
166
- - Use pretrained models (GROVER)
167
-
168
- ### Issue 3: Overfitting on Small Datasets
169
- **Problem**: Model memorizes training data.
170
- **Solutions**:
171
- - Use stronger regularization (increase dropout to 0.5)
172
- - Use simpler models (Random Forest instead of deep learning)
173
- - Apply transfer learning (ChemBERTa, GROVER)
174
- - Collect more data
175
-
176
- ### Issue 4: Import Errors
177
- **Problem**: `No module named 'torch'` / `No module named 'tensorflow'` warnings, or model classes fail to import.
178
- **Solution**: DeepChem loads lazily — install the backend that matches your model, then add the matching extra:
179
- ```bash
180
- uv pip install deepchem # loaders, featurizers, MoleculeNet only
181
- uv pip install 'deepchem[torch]' # GCN, GAT, AttentiveFP, HuggingFaceModel, GroverModel
182
- uv pip install 'deepchem[tensorflow]' # legacy Keras models
183
- uv pip install 'deepchem[jax]' # Haiku/JAX models
184
- ```
185
- Install PyTorch or TensorFlow with the correct CUDA build **before** the extra when using GPUs. Quote extras in zsh: `'deepchem[torch]'`.
186
-
187
- **Conda + PyTorch users:** If `import deepchem` fails with `undefined symbol: iJIT_NotifyEvent`, pin MKL below 2025 (`conda install "mkl<2025"`) — PyTorch wheels may be incompatible with MKL 2025.0.0.
188
-
189
- ## Reference Documentation
190
-
191
- This skill includes comprehensive reference documentation:
192
-
193
- ### `references/api_reference.md`
194
- Complete API documentation including:
195
- - All data loaders and their use cases
196
- - Dataset classes and when to use each
197
- - Complete featurizer catalog with selection guide
198
- - Model catalog organized by category (50+ models)
199
- - MoleculeNet dataset descriptions
200
- - Metrics and evaluation functions
201
- - Common code patterns
202
-
203
- **When to reference**: Search this file when you need specific API details, parameter names, or want to explore available options.
204
-
205
- ### `references/workflows.md`
206
- Eight detailed end-to-end workflows:
207
- 1. Molecular property prediction from SMILES
208
- 2. Using MoleculeNet benchmarks
209
- 3. Hyperparameter optimization
210
- 4. Transfer learning with pretrained models
211
- 5. Molecular generation with GANs
212
- 6. Materials property prediction
213
- 7. Protein sequence analysis
214
- 8. Custom model integration
215
-
216
- **When to reference**: Use these workflows as templates for implementing complete solutions.
217
-
218
- ## Installation
219
-
220
- Core package (data loaders, featurizers, MoleculeNet, scikit-learn wrappers):
221
-
222
- ```bash
223
- uv pip install deepchem
224
- ```
225
-
226
- Add the extra that matches your model backend (install PyTorch/TensorFlow/JAX first for GPU builds):
227
-
228
- ```bash
229
- uv pip install 'deepchem[torch]' # GNNs, TorchModel, HuggingFaceModel, GroverModel
230
- uv pip install 'deepchem[tensorflow]' # Keras/TensorFlow models
231
- uv pip install 'deepchem[jax]' # JAX/Haiku models
232
- uv pip install 'deepchem[dqc]' # Differentiable quantum chemistry (torch + xitorch)
233
- ```
234
-
235
- Nightly builds: `uv pip install --pre deepchem` (same extras apply with `--pre`).
236
-
237
- See [installation guide](https://deepchem.readthedocs.io/en/latest/get_started/installation.html) and [soft requirements](https://deepchem.readthedocs.io/en/latest/requirements.html) for optional dependencies per model class.
238
-
239
- ## Additional Resources
240
-
241
- - Official documentation: https://deepchem.readthedocs.io/
242
- - GitHub repository: https://github.com/deepchem/deepchem
243
- - Tutorials: https://deepchem.readthedocs.io/en/latest/get_started/tutorials.html
244
- - Paper: "MoleculeNet: A Benchmark for Molecular Machine Learning"
@@ -1,175 +0,0 @@
1
- ---
2
- name: deepspot-m
3
- description: Generate transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Use when you need spatial gene expression in log1p-CPM for 224x224 tiles at about 20x, want to query protein-coding genes by symbol instead of a fixed panel, or want to run prediction across a whole slide after tiling with histolab.
4
- license: PolyForm-Noncommercial-1.0.0
5
- compatibility: Needs deepspotm 1.0.0 from PyPI (Python 3.10 to 3.13) plus PyTorch. Weights at ratschlab/DeepSpotM on Hugging Face are gated and licensed CC-BY-NC-SA-4.0, so request access on the model page and then run huggingface-cli login. A CUDA GPU speeds up batched inference.
6
- allowed-tools: Read Write Edit Bash
7
- metadata:
8
- version: "1.0"
9
- skill-author: Ratschlab, ETH Zurich
10
- ---
11
-
12
- # DeepSpot-M
13
-
14
- ## Overview
15
-
16
- DeepSpot-M is a multimodal foundation model that maps a 224x224 H&E histology tile to
17
- spatial gene expression in log1p-CPM. The output is virtual spatial transcriptomics: one
18
- value per queried gene per tile, laid out on the grid the tiles came from.
19
-
20
- A LoRA-adapted pathology foundation backbone (Midnight) tokenises the tile. A
21
- cross-attention gene decoder lets each gene query attend to the patch tokens, and a gene
22
- router hypernetwork builds gene-specific projections from frozen biological embeddings
23
- (Evo 2, Orthrus, ProtT5, scGPT, Apertus). Genes enter the model as queryable embeddings
24
- rather than fixed output slots, so the released model covers a ~19k protein-coding gene
25
- panel including genes unseen in training. The panel ships with the weights as
26
- `tokens.csv` and is exposed as `model.gene_names`; genes outside it cannot be queried in
27
- this release.
28
-
29
- Applied to TCGA, the model produced a virtual spatial transcriptomics atlas of 28,664
30
- slides across 32 cancer types.
31
-
32
- ## Licensing
33
-
34
- The code is PolyForm Noncommercial 1.0.0 and the weights are CC-BY-NC-SA-4.0. Use it for
35
- noncommercial research and check both licences before redistributing outputs.
36
-
37
- ## Installation
38
-
39
- ```bash
40
- uv pip install deepspotm==1.0.0
41
- ```
42
-
43
- Version 1.0.0 targets Python 3.10 to 3.13 and pulls in PyTorch. Install the PyTorch build
44
- that matches your CUDA version first if you want GPU inference.
45
-
46
- ## Model access
47
-
48
- The weights are gated:
49
-
50
- 1. Open <https://huggingface.co/ratschlab/DeepSpotM> and request access.
51
- 2. Once access is granted, authenticate the machine that will download them:
52
-
53
- ```bash
54
- huggingface-cli login
55
- ```
56
-
57
- `from_pretrained` reads that cached token, so a login is needed once per machine.
58
-
59
- ## Quick start
60
-
61
- ```python
62
- from deepspotm import DeepSpotM
63
-
64
- model, image_processor = DeepSpotM.from_pretrained("ratschlab/DeepSpotM", source="scgpt")
65
-
66
- vals = model.predict_genes(image_processor(pil_tile).unsqueeze(0), ["EPCAM", "CD3D"])
67
- ```
68
-
69
- `pil_tile` is a PIL image of exactly 224x224 pixels. `image_processor` turns it into a
70
- tensor, `unsqueeze(0)` adds the batch dimension, and `predict_genes` takes the batch plus a
71
- list of HGNC gene symbols. Values come back in log1p-CPM, aligned with the gene list you
72
- passed, so keep that list beside the output to keep the columns labelled. Symbols must be
73
- in the released ~19k-gene panel (`model.gene_names`); an unknown symbol raises `KeyError`
74
- naming the offending genes.
75
-
76
- ## Tile requirements
77
-
78
- Tiles must be 224x224 RGB at roughly 20x magnification (about 0.5 microns per pixel). Check
79
- the size at the boundary of your pipeline rather than passing an unchecked crop through:
80
-
81
- ```python
82
- TILE_PX = 224
83
-
84
- def require_tile(tile):
85
- """Return an RGB 224x224 tile, or raise if the crop is the wrong size."""
86
- if tile.size != (TILE_PX, TILE_PX):
87
- raise ValueError(
88
- f"DeepSpot-M expects a {TILE_PX}x{TILE_PX} tile at about 20x "
89
- f"(~0.5 microns per pixel); got {tile.size[0]}x{tile.size[1]}. "
90
- "Re-tile at the matching level or resample the crop."
91
- )
92
- return tile.convert("RGB")
93
- ```
94
-
95
- Extract tiles at the slide level whose resolution is nearest 0.5 microns per pixel, then
96
- crop to 224x224 there. Resampling from a coarser level changes the texture the backbone
97
- reads.
98
-
99
- ## Keep the dependency optional
100
-
101
- `deepspotm` and its weights are a heavy, gated dependency. Import it inside the function
102
- that needs it so the surrounding project installs, imports and tests without it, and turn
103
- an `ImportError` into a message that names every step:
104
-
105
- ```python
106
- DEEPSPOTM_HELP = (
107
- "DeepSpot-M is unavailable. Install it with `uv pip install deepspotm==1.0.0`, request "
108
- "access to the gated weights at https://huggingface.co/ratschlab/DeepSpotM, then "
109
- "authenticate with `huggingface-cli login`."
110
- )
111
-
112
- def load_deepspotm(source="scgpt"):
113
- try:
114
- from deepspotm import DeepSpotM
115
- except ImportError as exc:
116
- raise RuntimeError(DEEPSPOTM_HELP) from exc
117
- return DeepSpotM.from_pretrained("ratschlab/DeepSpotM", source=source)
118
- ```
119
-
120
- ## Embedding sources
121
-
122
- `source` selects which frozen gene embedding the router builds projections from. It is one
123
- of five values:
124
-
125
- | `source` | Gene embedding |
126
- | --------- | --------------------------------- |
127
- | `evo2` | genomic sequence |
128
- | `orthrus` | RNA |
129
- | `prott5` | protein sequence |
130
- | `scgpt` | single-cell expression |
131
- | `apertus` | language model |
132
-
133
- Each gives a different view of gene identity. Pick one per run, and run the same tiles
134
- through more than one source when the choice matters to your analysis. See
135
- `references/api.md` for the full call surface, batching and device placement, gene symbol
136
- handling and output units.
137
-
138
- ## Whole slide workflow
139
-
140
- Prediction is per tile, so a slide-scale run is a tiling step followed by batched
141
- inference:
142
-
143
- 1. Extract 224x224 tiles on a grid with the `histolab` skill, keeping each tile's
144
- coordinates.
145
- 2. Process and stack tiles into batches with `torch.stack`.
146
- 3. Call `predict_genes` once per batch with the same gene list.
147
- 4. Concatenate the batches into a tiles-by-genes matrix and attach the coordinates.
148
-
149
- That matrix is the virtual spatial transcriptomics map for the slide, and it drops
150
- straight into `AnnData` for downstream spatial analysis. `references/whole_slide.md` has a
151
- worked loop, batch sizing and an `AnnData` assembly step.
152
-
153
- ## Common use cases
154
-
155
- - Spatial expression maps for marker genes across a tumour section.
156
- - Transcriptome-wide prediction over a slide cohort with no matching assay run.
157
- - Querying any of the ~19k panel genes by symbol, including genes unseen in training —
158
- far beyond the few hundred genes of a typical spatial assay panel.
159
- - Adding an expression channel to a morphology-only histology pipeline.
160
- - Building a slide-level cohort atlas, as done for TCGA.
161
-
162
- ## Detailed references
163
-
164
- - `references/api.md`: `from_pretrained` and `predict_genes` in full, the five embedding
165
- sources and how to choose, batching, device placement, gene symbol handling, and
166
- converting log1p-CPM output.
167
- - `references/whole_slide.md`: tiling with histolab, a slide-scale prediction loop,
168
- assembling and storing a tiles-by-genes matrix, and cohort-scale runs.
169
-
170
- ## Primary sources
171
-
172
- - Paper: <https://doi.org/10.64898/2026.06.19.26356060> (medRxiv, posted 22 June 2026)
173
- - Code: <https://github.com/ratschlab/DeepSpotM>
174
- - Weights: <https://huggingface.co/ratschlab/DeepSpotM>
175
- - PyPI: <https://pypi.org/project/deepspotm/>