@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,488 +0,0 @@
1
- ---
2
- name: diffdock
3
- description: DiffDock and DiffDock-L molecular docking. Use for protein-small-molecule pose prediction from PDB or sequence plus SMILES/SDF/MOL2, batch docking, virtual screening, and pose-confidence interpretation. Not for binding affinity prediction.
4
- allowed-tools: Read Write Edit Bash Glob Grep
5
- compatibility: Requires the DiffDock repository, Python 3.9 environment from upstream environment.yml or the official Docker image, RDKit, PyTorch/PyG, and optional CUDA GPU acceleration. Current guidance targets DiffDock v1.1.3 / DiffDock-L.
6
- license: MIT license
7
- metadata:
8
- version: "1.2"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # DiffDock: Molecular Docking with Diffusion Models
13
-
14
- ## Overview
15
-
16
- DiffDock is a diffusion-based deep learning tool for molecular docking that predicts 3D binding poses of small molecule ligands to protein targets. It represents the state-of-the-art in computational docking, crucial for structure-based drug discovery and chemical biology.
17
-
18
- **Core Capabilities:**
19
- - Predict ligand binding poses with high accuracy using deep learning
20
- - Support protein structures (PDB files) or sequences (via ESMFold)
21
- - Process single complexes or batch virtual screening campaigns
22
- - Generate confidence scores to assess prediction reliability
23
- - Handle diverse ligand inputs (SMILES, SDF, MOL2)
24
-
25
- **Key Distinction:** DiffDock predicts **binding poses** (3D structure) and **confidence** (prediction certainty), NOT binding affinity (ΔG, Kd). Always combine with scoring functions (GNINA, MM/GBSA) for affinity assessment.
26
-
27
- ## When to Use This Skill
28
-
29
- This skill should be used when:
30
-
31
- - "Dock this ligand to a protein" or "predict binding pose"
32
- - "Run molecular docking" or "perform protein-ligand docking"
33
- - "Virtual screening" or "screen compound library"
34
- - "Where does this molecule bind?" or "predict binding site"
35
- - Structure-based drug design or lead optimization tasks
36
- - Tasks involving PDB files + SMILES strings or ligand structures
37
- - Batch docking of multiple protein-ligand pairs
38
-
39
- ## Installation and Environment Setup
40
-
41
- ### Check Environment Status
42
-
43
- Before proceeding with DiffDock tasks, verify the environment setup:
44
-
45
- ```bash
46
- # Use the provided setup checker
47
- python scripts/setup_check.py
48
- ```
49
-
50
- This script validates Python version, PyTorch with CUDA, PyTorch Geometric, RDKit, ESM, and other dependencies.
51
-
52
- ### Installation Options
53
-
54
- **Option 1: Conda (Recommended)**
55
- ```bash
56
- git clone https://github.com/gcorso/DiffDock.git
57
- cd DiffDock
58
- conda env create --file environment.yml
59
- conda activate diffdock
60
- ```
61
-
62
- **Option 2: Docker**
63
- ```bash
64
- docker pull rbgcsail/diffdock
65
- docker run -it --gpus all --entrypoint /bin/bash rbgcsail/diffdock
66
- micromamba activate diffdock
67
- ```
68
-
69
- **Important Notes:**
70
- - GPU strongly recommended (10-100x speedup vs CPU)
71
- - First run pre-computes SO(2)/SO(3) lookup tables (~2-5 minutes)
72
- - Model checkpoints (~500MB) download automatically if not present
73
- - Current upstream release is DiffDock v1.1.3; DiffDock-L is the default model line in `default_inference_args.yaml`
74
-
75
- ## Core Workflows
76
-
77
- ### Workflow 1: Single Protein-Ligand Docking
78
-
79
- **Use Case:** Dock one ligand to one protein target
80
-
81
- **Input Requirements:**
82
- - Protein: PDB file OR amino acid sequence
83
- - Ligand: SMILES string OR structure file (SDF/MOL2)
84
-
85
- **Command:**
86
- ```bash
87
- python -m inference \
88
- --config default_inference_args.yaml \
89
- --protein_path protein.pdb \
90
- --ligand_description "CC(=O)Oc1ccccc1C(=O)O" \
91
- --out_dir results/single_docking/
92
- ```
93
-
94
- **Alternative (protein sequence):**
95
- ```bash
96
- python -m inference \
97
- --config default_inference_args.yaml \
98
- --protein_sequence "MSKGEELFTGVVPILVELDGDVNGHKF..." \
99
- --ligand_description ligand.sdf \
100
- --out_dir results/sequence_docking/
101
- ```
102
-
103
- **Output Structure:**
104
- ```
105
- results/single_docking/
106
- └── complex_0/
107
- ├── rank1.sdf # Convenience copy of top-ranked pose
108
- ├── rank1_confidence0.87.sdf # Top-ranked pose with confidence in filename
109
- ├── rank2_confidence0.42.sdf # Second-ranked pose
110
- ├── ...
111
- └── rank10_confidence-1.23.sdf # 10th pose (default: 10 samples)
112
- ```
113
-
114
- Current `inference.py` registers `--ligand_description` for single-complex runs. Some upstream README text still says `--ligand`; use `--ligand_description` unless your local checkout explicitly supports a `--ligand` alias.
115
-
116
- ### Workflow 2: Batch Processing Multiple Complexes
117
-
118
- **Use Case:** Dock multiple ligands to proteins, virtual screening campaigns
119
-
120
- **Step 1: Prepare Batch CSV**
121
-
122
- Use the provided script to create or validate batch input:
123
-
124
- ```bash
125
- # Create template
126
- python scripts/prepare_batch_csv.py --create --output batch_input.csv
127
-
128
- # Validate existing CSV
129
- python scripts/prepare_batch_csv.py my_input.csv --validate
130
- ```
131
-
132
- **CSV Format:**
133
- ```csv
134
- complex_name,protein_path,ligand_description,protein_sequence
135
- complex1,protein1.pdb,CC(=O)Oc1ccccc1C(=O)O,
136
- complex2,,COc1ccc(C#N)cc1,MSKGEELFT...
137
- complex3,protein3.pdb,ligand3.sdf,
138
- ```
139
-
140
- **Required Columns:**
141
- - `complex_name`: Unique identifier
142
- - `protein_path`: PDB file path (leave empty if using sequence)
143
- - `ligand_description`: SMILES string or ligand file path
144
- - `protein_sequence`: Amino acid sequence (leave empty if using PDB)
145
-
146
- **Step 2: Run Batch Docking**
147
-
148
- ```bash
149
- python -m inference \
150
- --config default_inference_args.yaml \
151
- --protein_ligand_csv batch_input.csv \
152
- --out_dir results/batch/ \
153
- --batch_size 10
154
- ```
155
-
156
- **For Large Virtual Screening (>100 compounds):**
157
-
158
- Pre-compute protein embeddings for faster processing:
159
- ```bash
160
- # Pre-compute embeddings
161
- python datasets/esm_embedding_preparation.py \
162
- --protein_ligand_csv screening_input.csv \
163
- --out_file protein_embeddings.pt
164
-
165
- # Run with pre-computed embeddings
166
- python -m inference \
167
- --config default_inference_args.yaml \
168
- --protein_ligand_csv screening_input.csv \
169
- --esm_embeddings_path protein_embeddings.pt \
170
- --out_dir results/screening/
171
- ```
172
-
173
- ### Workflow 3: Analyzing Results
174
-
175
- After docking completes, analyze confidence scores and rank predictions:
176
-
177
- ```bash
178
- # Analyze all results
179
- python scripts/analyze_results.py results/batch/
180
-
181
- # Show top 5 per complex
182
- python scripts/analyze_results.py results/batch/ --top 5
183
-
184
- # Filter by confidence threshold
185
- python scripts/analyze_results.py results/batch/ --threshold 0.0
186
-
187
- # Export to CSV
188
- python scripts/analyze_results.py results/batch/ --export summary.csv
189
-
190
- # Show top 20 predictions across all complexes
191
- python scripts/analyze_results.py results/batch/ --best 20
192
- ```
193
-
194
- The analysis script:
195
- - Parses confidence scores from all predictions
196
- - Classifies as High (>0), Moderate (-1.5 to 0), or Low (<-1.5)
197
- - Ranks predictions within and across complexes
198
- - Generates statistical summaries
199
- - Exports results to CSV for downstream analysis
200
-
201
- ## Confidence Score Interpretation
202
-
203
- **Understanding Scores:**
204
-
205
- | Score Range | Confidence Level | Interpretation |
206
- |------------|------------------|----------------|
207
- | **> 0** | High | Strong prediction, likely accurate |
208
- | **-1.5 to 0** | Moderate | Reasonable prediction, validate carefully |
209
- | **< -1.5** | Low | Uncertain prediction, requires validation |
210
-
211
- **Critical Notes:**
212
- 1. **Confidence ≠ Affinity**: High confidence means model certainty about structure, NOT strong binding
213
- 2. **Context Matters**: Adjust expectations for:
214
- - Large ligands (>500 Da): Lower confidence expected
215
- - Multiple protein chains: May decrease confidence
216
- - Novel protein families: May underperform
217
- 3. **Multiple Samples**: Review top 3-5 predictions, look for consensus
218
-
219
- **For detailed guidance:** Read `references/confidence_and_limitations.md` using the Read tool
220
-
221
- ## Parameter Customization
222
-
223
- ### Using Custom Configuration
224
-
225
- Create custom configuration for specific use cases:
226
-
227
- ```bash
228
- # Copy template
229
- cp assets/custom_inference_config.yaml my_config.yaml
230
-
231
- # Edit parameters (see template for presets)
232
- # Then run with custom config
233
- python -m inference \
234
- --config my_config.yaml \
235
- --protein_ligand_csv input.csv \
236
- --out_dir results/
237
- ```
238
-
239
- ### Key Parameters to Adjust
240
-
241
- **Sampling Density:**
242
- - `samples_per_complex: 10` → Increase to 20-40 for difficult cases
243
- - More samples = better coverage but longer runtime
244
-
245
- **Inference Steps:**
246
- - `inference_steps: 20` → Increase to 25-30 for higher accuracy
247
- - More steps = potentially better quality but slower
248
-
249
- **Temperature Parameters (control diversity):**
250
- - `temp_sampling_tor: 7.04` → Increase for flexible ligands (8-10)
251
- - `temp_sampling_tor: 7.04` → Decrease for rigid ligands (5-6)
252
- - Higher temperature = more diverse poses
253
-
254
- **Presets Available in Template:**
255
- 1. High Accuracy: More samples + steps, lower temperature
256
- 2. Fast Screening: Fewer samples, faster
257
- 3. Flexible Ligands: Increased torsion temperature
258
- 4. Rigid Ligands: Decreased torsion temperature
259
-
260
- **For complete parameter reference:** Read `references/parameters_reference.md` using the Read tool
261
-
262
- ## Advanced Techniques
263
-
264
- ### Ensemble Docking (Protein Flexibility)
265
-
266
- For proteins with known flexibility, dock to multiple conformations:
267
-
268
- ```python
269
- # Create ensemble CSV
270
- import pandas as pd
271
-
272
- conformations = ["conf1.pdb", "conf2.pdb", "conf3.pdb"]
273
- ligand = "CC(=O)Oc1ccccc1C(=O)O"
274
-
275
- data = {
276
- "complex_name": [f"ensemble_{i}" for i in range(len(conformations))],
277
- "protein_path": conformations,
278
- "ligand_description": [ligand] * len(conformations),
279
- "protein_sequence": [""] * len(conformations)
280
- }
281
-
282
- pd.DataFrame(data).to_csv("ensemble_input.csv", index=False)
283
- ```
284
-
285
- Run docking with increased sampling:
286
- ```bash
287
- python -m inference \
288
- --config default_inference_args.yaml \
289
- --protein_ligand_csv ensemble_input.csv \
290
- --samples_per_complex 20 \
291
- --out_dir results/ensemble/
292
- ```
293
-
294
- ### Integration with Scoring Functions
295
-
296
- DiffDock generates poses; combine with other tools for affinity:
297
-
298
- **GNINA (Fast neural network scoring):**
299
- ```bash
300
- for pose in results/single_docking/complex_0/*confidence*.sdf; do
301
- gnina -r protein.pdb -l "$pose" --score_only
302
- done
303
- ```
304
-
305
- **MM/GBSA (More accurate, slower):**
306
- Use AmberTools MMPBSA.py or gmx_MMPBSA after energy minimization
307
-
308
- **Free Energy Calculations (Most accurate):**
309
- Use OpenMM + OpenFE or GROMACS for FEP/TI calculations
310
-
311
- **Recommended Workflow:**
312
- 1. DiffDock → Generate poses with confidence scores
313
- 2. Visual inspection → Check structural plausibility
314
- 3. GNINA or MM/GBSA → Rescore and rank by affinity
315
- 4. Experimental validation → Biochemical assays
316
-
317
- ## Limitations and Scope
318
-
319
- **DiffDock IS Designed For:**
320
- - Small molecule ligands (typically 100-1000 Da)
321
- - Drug-like organic compounds
322
- - Small peptides (<20 residues)
323
- - Single or multi-chain proteins
324
-
325
- **DiffDock IS NOT Designed For:**
326
- - Large biomolecules (protein-protein docking) → Use DiffDock-PP or AlphaFold-Multimer
327
- - Large peptides (>20 residues) → Use alternative methods
328
- - Covalent docking → Use specialized covalent docking tools
329
- - Binding affinity prediction → Combine with scoring functions
330
- - Membrane proteins → Not specifically trained, use with caution
331
-
332
- **For complete limitations:** Read `references/confidence_and_limitations.md` using the Read tool
333
-
334
- ## Troubleshooting
335
-
336
- ### Common Issues
337
-
338
- **Issue: Low confidence scores across all predictions**
339
- - Cause: Large/unusual ligands, unclear binding site, protein flexibility
340
- - Solution: Increase `samples_per_complex` (20-40), try ensemble docking, validate protein structure
341
-
342
- **Issue: Out of memory errors**
343
- - Cause: GPU memory insufficient for batch size
344
- - Solution: Reduce `--batch_size 2` or process fewer complexes at once
345
-
346
- **Issue: Slow performance**
347
- - Cause: Running on CPU instead of GPU
348
- - Solution: Verify CUDA with `python -c "import torch; print(torch.cuda.is_available())"`, use GPU
349
-
350
- **Issue: Unrealistic binding poses**
351
- - Cause: Poor protein preparation, ligand too large, wrong binding site
352
- - Solution: Check protein for missing residues, remove far waters, consider specifying binding site
353
-
354
- **Issue: "Module not found" errors**
355
- - Cause: Missing dependencies or wrong environment
356
- - Solution: Run `python scripts/setup_check.py` to diagnose
357
-
358
- ### Performance Optimization
359
-
360
- **For Best Results:**
361
- 1. Use GPU (essential for practical use)
362
- 2. Pre-compute ESM embeddings for repeated protein use
363
- 3. Batch process multiple complexes together
364
- 4. Start with default parameters, then tune if needed
365
- 5. Validate protein structures (resolve missing residues)
366
- 6. Use canonical SMILES for ligands
367
-
368
- ## Graphical User Interface
369
-
370
- For interactive use, launch the web interface:
371
-
372
- ```bash
373
- python app/main.py
374
- # Navigate to http://localhost:7860
375
- ```
376
-
377
- Or use the online demo without installation:
378
- - https://huggingface.co/spaces/reginabarzilaygroup/DiffDock-Web
379
-
380
- ## Resources
381
-
382
- ### Helper Scripts (`scripts/`)
383
-
384
- **`prepare_batch_csv.py`**: Create and validate batch input CSV files
385
- - Create templates with example entries
386
- - Validate file paths and SMILES strings
387
- - Check for required columns and format issues
388
-
389
- **`analyze_results.py`**: Analyze confidence scores and rank predictions
390
- - Parse results from single or batch runs
391
- - Generate statistical summaries
392
- - Export to CSV for downstream analysis
393
- - Identify top predictions across complexes
394
-
395
- **`setup_check.py`**: Verify DiffDock environment setup
396
- - Check Python version and dependencies
397
- - Verify PyTorch and CUDA availability
398
- - Test RDKit and PyTorch Geometric installation
399
- - Provide installation instructions if needed
400
-
401
- ### Reference Documentation (`references/`)
402
-
403
- **`parameters_reference.md`**: Complete parameter documentation
404
- - All command-line options and configuration parameters
405
- - Default values and acceptable ranges
406
- - Temperature parameters for controlling diversity
407
- - Model checkpoint locations and version flags
408
-
409
- Read this file when users need:
410
- - Detailed parameter explanations
411
- - Fine-tuning guidance for specific systems
412
- - Alternative sampling strategies
413
-
414
- **`confidence_and_limitations.md`**: Confidence score interpretation and tool limitations
415
- - Detailed confidence score interpretation
416
- - When to trust predictions
417
- - Scope and limitations of DiffDock
418
- - Integration with complementary tools
419
- - Troubleshooting prediction quality
420
-
421
- Read this file when users need:
422
- - Help interpreting confidence scores
423
- - Understanding when NOT to use DiffDock
424
- - Guidance on combining with other tools
425
- - Validation strategies
426
-
427
- **`workflows_examples.md`**: Comprehensive workflow examples
428
- - Detailed installation instructions
429
- - Step-by-step examples for all workflows
430
- - Advanced integration patterns
431
- - Troubleshooting common issues
432
- - Best practices and optimization tips
433
-
434
- Read this file when users need:
435
- - Complete workflow examples with code
436
- - Integration with GNINA, OpenMM, or other tools
437
- - Virtual screening workflows
438
- - Ensemble docking procedures
439
-
440
- ### Assets (`assets/`)
441
-
442
- **`batch_template.csv`**: Template for batch processing
443
- - Pre-formatted CSV with required columns
444
- - Example entries showing different input types
445
- - Ready to customize with actual data
446
-
447
- **`custom_inference_config.yaml`**: Configuration template
448
- - Annotated YAML with all parameters
449
- - Four preset configurations for common use cases
450
- - Detailed comments explaining each parameter
451
- - Ready to customize and use
452
-
453
- ## Best Practices
454
-
455
- 1. **Always verify environment** with `setup_check.py` before starting large jobs
456
- 2. **Validate batch CSVs** with `prepare_batch_csv.py` to catch errors early
457
- 3. **Start with defaults** then tune parameters based on system-specific needs
458
- 4. **Generate multiple samples** (10-40) for robust predictions
459
- 5. **Visual inspection** of top poses before downstream analysis
460
- 6. **Combine with scoring** functions for affinity assessment
461
- 7. **Use confidence scores** for initial ranking, not final decisions
462
- 8. **Pre-compute embeddings** for virtual screening campaigns
463
- 9. **Document parameters** used for reproducibility
464
- 10. **Validate results** experimentally when possible
465
-
466
- ## Citations
467
-
468
- When using DiffDock, cite the appropriate papers:
469
-
470
- **DiffDock-L (current default model):**
471
- ```
472
- Corso et al. (2024) "Deep Confident Steps to New Pockets: Strategies for Docking Generalization"
473
- ICLR 2024, arXiv:2402.18396
474
- ```
475
-
476
- **Original DiffDock:**
477
- ```
478
- Corso et al. (2023) "DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking"
479
- ICLR 2023, arXiv:2210.01776
480
- ```
481
-
482
- ## Additional Resources
483
-
484
- - **GitHub Repository**: https://github.com/gcorso/DiffDock
485
- - **Online Demo**: https://huggingface.co/spaces/reginabarzilaygroup/DiffDock-Web
486
- - **DiffDock-L Paper**: https://arxiv.org/abs/2402.18396
487
- - **Original Paper**: https://arxiv.org/abs/2210.01776
488
-