@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,348 +0,0 @@
1
- ---
2
- name: molfeat
3
- description: Molecular featurization for ML (100+ featurizers). ECFP, MACCS, descriptors, pretrained models (ChemBERTa), convert SMILES to features, for QSAR and molecular ML.
4
- license: Apache-2.0 license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.9–3.10 (molfeat 0.11.0 does not support 3.11+). Requires datamol, PyTorch, and optional extras for GNN/transformer models.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # Molfeat - Molecular Featurization Hub
13
-
14
- ## Overview
15
-
16
- Molfeat is a comprehensive Python library for molecular featurization that unifies 100+ pre-trained embeddings and hand-crafted featurizers. Convert chemical structures (SMILES strings or RDKit molecules) into numerical representations for machine learning tasks including QSAR modeling, virtual screening, similarity searching, and deep learning applications. Features fast parallel processing, scikit-learn compatible transformers, and built-in caching.
17
-
18
- **Version note:** Examples target **molfeat 0.11.0** (PyPI stable, May 2025). Requires **Python 3.9–3.10** (`requires-python` caps below 3.11). Depends on **datamol ≥0.8.0** and **PyTorch ≥1.13**. Since 0.8.7, prefer datamol `Mol` objects over raw `rdkit.Chem.Mol`. Since 0.10.1, fingerprint calculators use RDKit's `rdFingerprintGenerator` API internally. Since 0.11.0, pretrained models load in memory and base models are set to PyTorch evaluation mode automatically.
19
-
20
- ## When to Use This Skill
21
-
22
- This skill should be used when working with:
23
- - **Molecular machine learning**: Building QSAR/QSPR models, property prediction
24
- - **Virtual screening**: Ranking compound libraries for biological activity
25
- - **Similarity searching**: Finding structurally similar molecules
26
- - **Chemical space analysis**: Clustering, visualization, dimensionality reduction
27
- - **Deep learning**: Training neural networks on molecular data
28
- - **Featurization pipelines**: Converting SMILES to ML-ready representations
29
- - **Cheminformatics**: Any task requiring molecular feature extraction
30
-
31
- ## Installation
32
-
33
- Use a Python 3.9 or 3.10 environment (molfeat does not install on 3.11+ as of 0.11.0):
34
-
35
- ```bash
36
- uv pip install "molfeat==0.11.0"
37
-
38
- # With all pip-installable optional dependencies
39
- uv pip install "molfeat[all]==0.11.0"
40
- ```
41
-
42
- **Optional dependency extras (PyPI):**
43
- - `molfeat[dgl]` — GNN models (GIN variants); upstream recommends `dgl<=2.0` (graphbolt issues in newer DGL)
44
- - `molfeat[graphormer]` — Graphormer models
45
- - `molfeat[transformer]` — ChemBERTa, ChemGPT, MolT5
46
- - `molfeat[fcd]` — FCD descriptors
47
- - `molfeat[pyg]` — PyTorch Geometric featurizers
48
- - `molfeat[viz]` — NGLView visualization widgets
49
-
50
- **External featurizers:** MAP4 is not bundled in molfeat extras — install from [reymond-group/map4](https://github.com/reymond-group/map4) separately. Some heavy deps (DGL, dgllife, graphormer-pretrained) are easier via conda-forge; see [optional dependencies](https://molfeat-docs.datamol.io/stable/).
51
-
52
- ## Core Concepts
53
-
54
- Molfeat organizes featurization into three hierarchical classes:
55
-
56
- ### 1. Calculators (`molfeat.calc`)
57
-
58
- Callable objects that convert individual molecules into feature vectors. Accept RDKit `Chem.Mol` objects or SMILES strings.
59
-
60
- **Use calculators for:**
61
- - Single molecule featurization
62
- - Custom processing loops
63
- - Direct feature computation
64
-
65
- **Example:**
66
- ```python
67
- from molfeat.calc import FPCalculator
68
-
69
- calc = FPCalculator("ecfp", radius=3, fpSize=2048)
70
- features = calc("CCO") # Returns numpy array (2048,)
71
- ```
72
-
73
- ### 2. Transformers (`molfeat.trans`)
74
-
75
- Scikit-learn compatible transformers that wrap calculators for batch processing with parallelization.
76
-
77
- **Use transformers for:**
78
- - Batch featurization of molecular datasets
79
- - Integration with scikit-learn pipelines
80
- - Parallel processing (automatic CPU utilization)
81
-
82
- **Example:**
83
- ```python
84
- from molfeat.trans import MoleculeTransformer
85
- from molfeat.calc import FPCalculator
86
-
87
- transformer = MoleculeTransformer(FPCalculator("ecfp"), n_jobs=-1)
88
- features = transformer(smiles_list) # Parallel processing
89
- ```
90
-
91
- ### 3. Pretrained Transformers (`molfeat.trans.pretrained`)
92
-
93
- Specialized transformers for deep learning models with batched inference and caching.
94
-
95
- **Use pretrained transformers for:**
96
- - State-of-the-art molecular embeddings
97
- - Transfer learning from large chemical datasets
98
- - Deep learning feature extraction
99
-
100
- **Example:**
101
- ```python
102
- from molfeat.trans.pretrained import PretrainedMolTransformer
103
-
104
- transformer = PretrainedMolTransformer("ChemBERTa-77M-MLM", n_jobs=-1)
105
- embeddings = transformer(smiles_list) # Deep learning embeddings
106
- ```
107
-
108
- ## Quick Start Workflow
109
-
110
- ### Basic Featurization
111
-
112
- ```python
113
- import datamol as dm
114
- from molfeat.calc import FPCalculator
115
- from molfeat.trans import MoleculeTransformer
116
-
117
- # Load molecular data
118
- smiles = ["CCO", "CC(=O)O", "c1ccccc1", "CC(C)O"]
119
-
120
- # Create calculator and transformer
121
- calc = FPCalculator("ecfp", radius=3)
122
- transformer = MoleculeTransformer(calc, n_jobs=-1)
123
-
124
- # Featurize molecules
125
- features = transformer(smiles)
126
- print(f"Shape: {features.shape}") # (4, 2048)
127
- ```
128
-
129
- ### Save and Load Configuration
130
-
131
- ```python
132
- # Save featurizer configuration for reproducibility
133
- transformer.to_state_yaml_file("featurizer_config.yml")
134
-
135
- # Reload exact configuration
136
- loaded = MoleculeTransformer.from_state_yaml_file("featurizer_config.yml")
137
- ```
138
-
139
- ### Handle Errors Gracefully
140
-
141
- ```python
142
- # Process dataset with potentially invalid SMILES
143
- transformer = MoleculeTransformer(
144
- calc,
145
- n_jobs=-1,
146
- ignore_errors=True, # Continue on failures
147
- verbose=True # Log error details
148
- )
149
-
150
- features = transformer(smiles_with_errors)
151
- # Returns None for failed molecules
152
- ```
153
-
154
- ## Choosing a Featurizer and Common Workflows
155
-
156
- Featurizer choice by task — traditional ML (RF, SVM, XGBoost), deep learning, similarity
157
- searching, and pharmacophore-based approaches — plus worked workflows for QSAR model
158
- building, virtual screening, similarity search, scikit-learn pipeline integration, and
159
- comparing multiple featurizers, are in
160
- [references/choosing_a_featurizer.md](references/choosing_a_featurizer.md).
161
-
162
- The full featurizer list is in
163
- [references/available_featurizers.md](references/available_featurizers.md); more examples
164
- are in [references/examples.md](references/examples.md).
165
-
166
- ## Discovering Available Featurizers
167
-
168
- Use the ModelStore to explore all available featurizers:
169
-
170
- ```python
171
- from molfeat.store.modelstore import ModelStore
172
-
173
- store = ModelStore()
174
-
175
- # List all available models
176
- all_models = store.available_models
177
- print(f"Total featurizers: {len(all_models)}")
178
-
179
- # Search for specific models
180
- chemberta_models = store.search(name="ChemBERTa")
181
- for model in chemberta_models:
182
- print(f"- {model.name}: {model.description}")
183
-
184
- # Get usage information
185
- model_card = store.search(name="ChemBERTa-77M-MLM")[0]
186
- model_card.usage() # Display usage examples
187
-
188
- # Load model
189
- transformer = store.load("ChemBERTa-77M-MLM")
190
- ```
191
-
192
- ## Advanced Features
193
-
194
- ### Custom Preprocessing
195
-
196
- ```python
197
- class CustomTransformer(MoleculeTransformer):
198
- def preprocess(self, mol):
199
- """Custom preprocessing pipeline"""
200
- if isinstance(mol, str):
201
- mol = dm.to_mol(mol)
202
- mol = dm.standardize_mol(mol)
203
- mol = dm.remove_salts(mol)
204
- return mol
205
-
206
- transformer = CustomTransformer(FPCalculator("ecfp"), n_jobs=-1)
207
- ```
208
-
209
- ### Batch Processing Large Datasets
210
-
211
- ```python
212
- import numpy as np
213
-
214
- def featurize_in_chunks(smiles_list, transformer, chunk_size=10000):
215
- """Process large datasets in chunks to manage memory"""
216
- all_features = []
217
- for i in range(0, len(smiles_list), chunk_size):
218
- chunk = smiles_list[i:i+chunk_size]
219
- features = transformer(chunk)
220
- all_features.append(features)
221
- return np.vstack(all_features)
222
- ```
223
-
224
- ### Caching Expensive Embeddings
225
-
226
- Prefer molfeat's built-in pretrained-model cache when possible. For custom embedding caches, use NumPy arrays instead of pickle (pickle can execute arbitrary code when loading untrusted files):
227
-
228
- ```python
229
- import numpy as np
230
- from pathlib import Path
231
-
232
- cache_file = Path("embeddings_cache.npz") # fixed path under your project
233
- transformer = PretrainedMolTransformer("ChemBERTa-77M-MLM", n_jobs=-1)
234
-
235
- if cache_file.exists():
236
- embeddings = np.load(cache_file)["embeddings"]
237
- else:
238
- embeddings = transformer(smiles_list)
239
- np.savez(cache_file, embeddings=embeddings)
240
- ```
241
-
242
- ## Performance Tips
243
-
244
- 1. **Use parallelization**: Set `n_jobs=-1` to utilize all CPU cores
245
- 2. **Batch processing**: Process multiple molecules at once instead of loops
246
- 3. **Choose appropriate featurizers**: Fingerprints are faster than deep learning models
247
- 4. **Cache pretrained models**: Leverage built-in caching for repeated use
248
- 5. **Use float32**: Set `dtype=np.float32` when precision allows
249
- 6. **Handle errors efficiently**: Use `ignore_errors=True` for large datasets
250
-
251
- ## Common Featurizers Reference
252
-
253
- **Quick reference for frequently used featurizers:**
254
-
255
- | Featurizer | Type | Dimensions | Speed | Use Case |
256
- |------------|------|------------|-------|----------|
257
- | `ecfp` | Fingerprint | 2048 | Fast | General purpose |
258
- | `maccs` | Fingerprint | 167 | Very fast | Scaffold similarity |
259
- | `desc2D` | Descriptors | 200+ | Fast | Interpretable models |
260
- | `mordred` | Descriptors | 1800+ | Medium | Comprehensive features |
261
- | `map4` | Fingerprint | 1024 | Fast | Large-scale screening |
262
- | `ChemBERTa-77M-MLM` | Deep learning | 768 | Slow* | Transfer learning |
263
- | `gin-supervised-masking` | GNN | Variable | Slow* | Graph-based models |
264
-
265
- *First run is slow; subsequent runs benefit from caching
266
-
267
- ## Resources
268
-
269
- This skill includes comprehensive reference documentation:
270
-
271
- ### references/api_reference.md
272
- Complete API documentation covering:
273
- - `molfeat.calc` - All calculator classes and parameters
274
- - `molfeat.trans` - Transformer classes and methods
275
- - `molfeat.store` - ModelStore usage
276
- - Common patterns and integration examples
277
- - Performance optimization tips
278
-
279
- **When to load:** Reference when implementing specific calculators, understanding transformer parameters, or integrating with scikit-learn/PyTorch.
280
-
281
- ### references/available_featurizers.md
282
- Comprehensive catalog of all 100+ featurizers organized by category:
283
- - Transformer-based language models (ChemBERTa, ChemGPT)
284
- - Graph neural networks (GIN, Graphormer)
285
- - Molecular descriptors (RDKit, Mordred)
286
- - Fingerprints (ECFP, MACCS, MAP4, and 15+ others)
287
- - Pharmacophore descriptors (CATS, Gobbi)
288
- - Shape descriptors (USR, ElectroShape)
289
- - Scaffold-based descriptors
290
-
291
- **When to load:** Reference when selecting the optimal featurizer for a specific task, exploring available options, or understanding featurizer characteristics.
292
-
293
- **Search tip:** Use grep to find specific featurizer types:
294
- ```bash
295
- grep -i "chembert" references/available_featurizers.md
296
- grep -i "pharmacophore" references/available_featurizers.md
297
- ```
298
-
299
- ### references/examples.md
300
- Practical code examples for common scenarios:
301
- - Installation and quick start
302
- - Calculator and transformer examples
303
- - Pretrained model usage
304
- - Scikit-learn and PyTorch integration
305
- - Virtual screening workflows
306
- - QSAR model building
307
- - Similarity searching
308
- - Troubleshooting and best practices
309
-
310
- **When to load:** Reference when implementing specific workflows, troubleshooting issues, or learning molfeat patterns.
311
-
312
- ## Troubleshooting
313
-
314
- ### Invalid Molecules
315
- Enable error handling to skip invalid SMILES:
316
- ```python
317
- transformer = MoleculeTransformer(
318
- calc,
319
- ignore_errors=True,
320
- verbose=True
321
- )
322
- ```
323
-
324
- ### Memory Issues with Large Datasets
325
- Process in chunks or use streaming approaches for datasets > 100K molecules.
326
-
327
- ### Pretrained Model Dependencies
328
- Some models require additional packages. Install specific extras (pin version for reproducibility):
329
- ```bash
330
- uv pip install "molfeat[transformer]==0.11.0" # For ChemBERTa/ChemGPT
331
- uv pip install "molfeat[dgl]==0.11.0" # For GIN models
332
- uv pip install "molfeat[graphormer]==0.11.0" # For Graphormer
333
- ```
334
-
335
- ### Reproducibility
336
- Save exact configurations and document versions:
337
- ```python
338
- transformer.to_state_yaml_file("config.yml")
339
- import molfeat
340
- print(f"molfeat version: {molfeat.__version__}")
341
- ```
342
-
343
- ## Additional Resources
344
-
345
- - **Official Documentation**: https://molfeat-docs.datamol.io/
346
- - **GitHub Repository**: https://github.com/datamol-io/molfeat
347
- - **PyPI Package**: https://pypi.org/project/molfeat/
348
- - **Tutorial**: https://portal.valencelabs.com/datamol/post/types-of-featurizers-b1e8HHrbFMkbun6
@@ -1,178 +0,0 @@
1
- ---
2
- name: ncats-arax
3
- description: Queries the NCATS Translator ARAX production API for bounded, typed, provenance-rich one-hop and endpoint-pinned two-hop biomedical knowledge-graph relationships. Use for Biolink-constrained RTX-KG2 lookup, explicit selected-provider ARAX federation, separate entity normalization, qualifier-aware graph traversal, and inspection of TRAPI edge bindings, publications, and knowledge-source provenance. Do not use for inference, ranking, open-ended pathfinding, clinical guidance, or sensitive queries.
4
- allowed-tools: Read Bash
5
- license: MIT
6
- compatibility: Requires Python 3.10+ and outbound HTTPS access to arax.transltr.io. The client uses only the Python standard library and needs no API key. Queries and caller metadata may be publicly visible; never submit sensitive or patient-specific content.
7
- metadata:
8
- version: "1.0"
9
- skill-author: neuroepithelial
10
- ---
11
-
12
- # NCATS ARAX
13
-
14
- Use ARAX as a constrained knowledge-graph lookup service. Submit reviewed CURIEs and explicit
15
- Biolink types, preserve the exact TRAPI exchange, inspect query-edge bindings and provenance, and
16
- treat every returned path as a candidate for subsequent verification.
17
-
18
- Read [query-contract.md](references/query-contract.md) before constructing a query. Read
19
- [output-schema.md](references/output-schema.md) when interpreting saved artifacts, warnings,
20
- provenance, or partial results.
21
-
22
- ## Safety boundary
23
-
24
- - Use only public, nonsensitive research questions. ARAX status facilities may expose query and
25
- caller metadata even when `store=false` is requested.
26
- - Do not submit patient information, confidential research questions, unpublished compound
27
- programs, or proprietary target hypotheses.
28
- - Do not present a returned path as a validated mechanism or clinical recommendation.
29
- - Report a zero as "not returned under these constraints," never as evidence that no relationship
30
- exists.
31
- - Describe position as unscored response order, never rank.
32
- - Verify important candidates with literature and authoritative databases separately.
33
-
34
- ## Workflow
35
-
36
- 1. Normalize free text separately, then review and report the proposed CURIE and category.
37
- 2. Choose a typed one-hop query or an exactly two-hop query with both endpoints pinned.
38
- 3. Use default RTX-KG2 lookup unless the user explicitly names two to five providers.
39
- 4. Acknowledge that the biomedical query is public and choose a new or empty output directory.
40
- 5. Run the client once. Do not silently change provider selection or expansion order after a
41
- failure or empty result.
42
- 6. Inspect `summary.json` for bounded bindings and provenance and `response.json` for the exact
43
- TRAPI payload.
44
- 7. Verify scientifically important paths outside ARAX.
45
-
46
- ## Preflight
47
-
48
- Check the production OpenAPI without making a biomedical query:
49
-
50
- ```bash
51
- python skills/ncats-arax/scripts/arax_client.py preflight
52
- ```
53
-
54
- The client verifies that the service identifies itself as ARAX, exposes `/query`, and reports a
55
- supported TRAPI version. A nonproduction endpoint or untested TRAPI series requires an explicit
56
- override; neither override changes the fixed query shapes or operations.
57
-
58
- ## Normalize an entity
59
-
60
- Normalization is review-only and never triggers a graph query:
61
-
62
- ```bash
63
- python skills/ncats-arax/scripts/arax_client.py normalize "primary myelofibrosis" \
64
- --expected-category biolink:Disease \
65
- --max-synonyms 10 \
66
- --acknowledge-public-query \
67
- --output-dir outputs/normalize-myelofibrosis
68
- ```
69
-
70
- Review the canonical identifier, name, category, and synonym preview before using a CURIE. Report
71
- all CURIEs and categories regardless of query outcome. A category warning or zero result is a
72
- reason to curate the identifier, not to chain automatically to `/query`.
73
-
74
- ## One-hop lookup
75
-
76
- Pin at least one endpoint and type both nodes:
77
-
78
- ```bash
79
- python skills/ncats-arax/scripts/arax_client.py one-hop \
80
- --subject-id CHEBI:31690 \
81
- --subject-category biolink:SmallMolecule \
82
- --predicate biolink:affects \
83
- --object-id NCBIGene:25 \
84
- --object-category biolink:Gene \
85
- --qualifier biolink:object_aspect_qualifier=activity_or_abundance \
86
- --qualifier biolink:object_direction_qualifier=decreased \
87
- --acknowledge-public-query \
88
- --output-dir outputs/imatinib-abl1
89
- ```
90
-
91
- Lookup mode is the default and fixes expansion to `infores:rtx-kg2`. It defaults to 20 results.
92
- Use `--result-limit N` to request 1-50 results; 50 is the hard cap in either mode.
93
-
94
- ## Endpoint-pinned two-hop lookup
95
-
96
- Use exactly one typed, unpinned intermediate node:
97
-
98
- ```bash
99
- python skills/ncats-arax/scripts/arax_client.py two-hop \
100
- --subject-id CHEBI:66901 \
101
- --subject-category biolink:SmallMolecule \
102
- --predicate-1 biolink:affects \
103
- --intermediate-category biolink:Gene \
104
- --predicate-2 biolink:associated_with \
105
- --object-id MONDO:0009061 \
106
- --object-category biolink:Disease \
107
- --qualifier-1 biolink:object_aspect_qualifier=activity_or_abundance \
108
- --qualifier-1 biolink:object_direction_qualifier=increased \
109
- --expand-order right-first \
110
- --acknowledge-public-query \
111
- --output-dir outputs/ivacaftor-cystic-fibrosis
112
- ```
113
-
114
- Right-first expansion is the default. If an empty result merits another attempt, run a new query
115
- explicitly with `--expand-order left-first` and keep the runs separate.
116
-
117
- ## Selected-provider federation
118
-
119
- Federation is explicit and accepts two to five named providers:
120
-
121
- ```bash
122
- python skills/ncats-arax/scripts/arax_client.py one-hop \
123
- --subject-id CHEBI:31690 \
124
- --subject-category biolink:SmallMolecule \
125
- --predicate biolink:affects \
126
- --object-id NCBIGene:25 \
127
- --object-category biolink:Gene \
128
- --mode federated \
129
- --kp infores:rtx-kg2 \
130
- --kp infores:molepro \
131
- --acknowledge-public-query \
132
- --output-dir outputs/federated-imatinib-abl1
133
- ```
134
-
135
- Federation defaults to the hard maximum of 50 results. Provider errors may coexist with useful
136
- results; such a run exits 7 after retaining its artifacts and is marked partial.
137
-
138
- ## Inspect saved provenance
139
-
140
- Rebuild a bounded summary without network access:
141
-
142
- ```bash
143
- python skills/ncats-arax/scripts/arax_client.py summarize \
144
- --request outputs/ivacaftor-cystic-fibrosis/request.json \
145
- --response outputs/ivacaftor-cystic-fibrosis/response.json \
146
- --format text
147
- ```
148
-
149
- The inspector accepts only the same constrained request shapes and fixed operations that the live
150
- commands generate. Use `--format json` for the normalized view on standard output.
151
-
152
- ## Interpret results
153
-
154
- - Follow each analysis's query-edge bindings; do not summarize every knowledge-graph edge.
155
- - Preserve the physical edge subject, predicate, object, and qualifier values returned by ARAX.
156
- Returned predicates or qualifier aspects may be more specific than the query constraint.
157
- - Inspect all source objects, including primary, aggregator, supporting-data, upstream-resource,
158
- and source-record URL fields.
159
- - Treat `publication_availability: not_returned` as missing metadata, not evidence that no
160
- publications exist.
161
- - Treat missing auxiliary-graph references and provider failures as explicit warnings.
162
- - Consult the raw response whenever the bounded summary omits detail or the service response is
163
- partial, unfamiliar, or scientifically surprising.
164
-
165
- ## Deliberate exclusions
166
-
167
- The client has no raw-query, workflow, operation, overlay, ranking, inference, link-prediction,
168
- Pathfinder, ARS, batch, all-provider, three-hop, cache, daemon, SDK, MCP, or
169
- natural-language-to-TRAPI surface. Do not work around those limits with direct HTTP calls under
170
- this skill.
171
-
172
- ## Official references
173
-
174
- - [ARAX documentation](https://ncatstranslator.github.io/TranslatorTechnicalDocumentation/architecture/ara/arax/)
175
- - [ARAX production OpenAPI](https://arax.transltr.io/api/arax/v1.4/openapi.json)
176
- - [ARAXi operation documentation](https://github.com/RTXteam/RTX/blob/master/code/ARAX/Documentation/DSL_Documentation.md)
177
- - [Translator Reasoner API](https://github.com/NCATSTranslator/ReasonerAPI)
178
- - [Biolink Model](https://biolink.github.io/biolink-model/)