@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,399 +0,0 @@
1
- ---
2
- name: bioservices
3
- description: Unified Python interface to 40+ bioinformatics services. Use when querying multiple databases (UniProt, KEGG, ChEMBL, Reactome) in a single workflow with consistent API. Best for cross-database analysis, ID mapping across services. For quick single-database lookups use gget; for sequence/file manipulation use biopython.
4
- license: GPLv3 license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.9–3.12 and internet access to 40+ bioinformatics web APIs. NCBI BLAST requires a contact email (`NCBI_EMAIL` env var or explicit parameter).
7
- metadata:
8
- version: "1.3"
9
- skill-author: K-Dense Inc.
10
- openclaw:
11
- envVars:
12
- - name: NCBI_EMAIL
13
- required: false
14
- description: Email for NCBI service identification.
15
- ---
16
-
17
- # BioServices
18
-
19
- ## Overview
20
-
21
- BioServices is a Python package providing programmatic access to approximately 40 bioinformatics web services and databases. Retrieve biological data, perform cross-database queries, map identifiers, analyze sequences, and integrate multiple biological resources in Python workflows. The package handles both REST and SOAP/WSDL protocols transparently.
22
-
23
- **Version note:** Examples target **bioservices 1.16.0** (PyPI, Mar 2026). Requires **Python 3.9–3.12**. UniProt REST changes in mid-2022 (bioservices ≥1.10) mainly affect tabular `columns` names — see upstream `_legacy_names` if parsing breaks. ChEMBL wrappers changed at 1.6.0 (2018 API); use `get_similarity`, `get_substructure`, `get_molecule` instead of pre-1.6 method names.
24
-
25
- ## When to Use This Skill
26
-
27
- This skill should be used when:
28
- - Retrieving protein sequences, annotations, or structures from UniProt, PDB, Pfam
29
- - Analyzing metabolic pathways and gene functions via KEGG or Reactome
30
- - Searching compound databases (ChEBI, ChEMBL, PubChem) for chemical information
31
- - Converting identifiers between different biological databases (KEGG↔UniProt, compound IDs)
32
- - Running sequence similarity searches (BLAST, MUSCLE alignment)
33
- - Querying gene ontology terms (QuickGO, GO annotations)
34
- - Accessing protein-protein interaction data (PSICQUIC, IntactComplex)
35
- - Mining genomic data (BioMart, ArrayExpress, ENA)
36
- - Integrating data from multiple bioinformatics resources in a single workflow
37
-
38
- ## Core Capabilities
39
-
40
- ### 1. Protein Analysis
41
-
42
- Retrieve protein information, sequences, and functional annotations:
43
-
44
- ```python
45
- from bioservices import UniProt
46
-
47
- u = UniProt(verbose=False)
48
-
49
- # Search for protein by name
50
- results = u.search("ZAP70_HUMAN", frmt="tab", columns="id,genes,organism")
51
-
52
- # Retrieve FASTA sequence
53
- sequence = u.retrieve("P43403", "fasta")
54
-
55
- # Map identifiers between databases
56
- kegg_ids = u.mapping(fr="UniProtKB_AC-ID", to="KEGG", query="P43403")
57
- ```
58
-
59
- **Key methods:**
60
- - `search()`: Query UniProt with flexible search terms
61
- - `retrieve()`: Get protein entries in various formats (FASTA, XML, tab)
62
- - `mapping()`: Convert identifiers between databases
63
-
64
- Reference: `references/services_reference.md` for complete UniProt API details.
65
-
66
- ### 2. Pathway Discovery and Analysis
67
-
68
- Access KEGG pathway information for genes and organisms:
69
-
70
- ```python
71
- from bioservices import KEGG
72
-
73
- k = KEGG()
74
- k.organism = "hsa" # Set to human
75
-
76
- # Search for organisms
77
- k.lookfor_organism("droso") # Find Drosophila species
78
-
79
- # Find pathways by name
80
- k.lookfor_pathway("B cell") # Returns matching pathway IDs
81
-
82
- # Get pathways containing specific genes
83
- pathways = k.get_pathway_by_gene("7535", "hsa") # ZAP70 gene
84
-
85
- # Retrieve and parse pathway data
86
- data = k.get("hsa04660")
87
- parsed = k.parse(data)
88
-
89
- # Extract pathway interactions
90
- interactions = k.parse_kgml_pathway("hsa04660")
91
- relations = interactions['relations'] # Protein-protein interactions
92
-
93
- # Convert to Simple Interaction Format
94
- sif_data = k.pathway2sif("hsa04660")
95
- ```
96
-
97
- **Key methods:**
98
- - `lookfor_organism()`, `lookfor_pathway()`: Search by name
99
- - `get_pathway_by_gene()`: Find pathways containing genes
100
- - `parse_kgml_pathway()`: Extract structured pathway data
101
- - `pathway2sif()`: Get protein interaction networks
102
-
103
- Reference: `references/workflow_patterns.md` for complete pathway analysis workflows.
104
-
105
- ### 3. Compound Database Searches
106
-
107
- Search and cross-reference compounds across multiple databases:
108
-
109
- ```python
110
- from bioservices import KEGG, UniChem
111
-
112
- k = KEGG()
113
-
114
- # Search compounds by name
115
- results = k.find("compound", "Geldanamycin") # Returns cpd:C11222
116
-
117
- # Get compound information with database links
118
- compound_info = k.get("cpd:C11222") # Includes ChEBI links
119
-
120
- # Cross-reference KEGG → ChEMBL using UniChem
121
- u = UniChem()
122
- chembl_id = u.get_compound_id_from_kegg("C11222") # Returns CHEMBL278315
123
- ```
124
-
125
- **Version caveat:** the per-source `get_compound_id_from_*` helpers are gone from
126
- bioservices 1.16.0 — check `hasattr(u, "get_compound_id_from_kegg")` first, and
127
- otherwise use the current UniChem API (`u.get_compounds(compound, source_type)`
128
- and read `res["compounds"][0]["sources"]`). ChEMBL lookups follow the same rule:
129
- `get_molecule`, not the pre-1.6 `get_compound_by_chemblId`.
130
-
131
- **Common workflow:**
132
- 1. Search compound by name in KEGG
133
- 2. Extract KEGG compound ID
134
- 3. Use UniChem for KEGG → ChEMBL mapping
135
- 4. ChEBI IDs are often provided in KEGG entries
136
-
137
- Reference: `references/identifier_mapping.md` for complete cross-database mapping guide.
138
-
139
- ### 4. Sequence Analysis
140
-
141
- Run BLAST searches and sequence alignments. NCBI requires a contact email — prefer the `NCBI_EMAIL` environment variable (same convention as BioPython Entrez and other repo skills):
142
-
143
- ```python
144
- import os
145
- from bioservices import NCBIblast
146
-
147
- s = NCBIblast(verbose=False)
148
- email = os.environ["NCBI_EMAIL"] # set before running: export NCBI_EMAIL=you@lab.org
149
-
150
- # Run BLASTP against UniProtKB
151
- jobid = s.run(
152
- program="blastp",
153
- sequence=protein_sequence,
154
- stype="protein",
155
- database="uniprotkb",
156
- email=email,
157
- )
158
-
159
- # Check job status and retrieve results
160
- s.getStatus(jobid)
161
- results = s.getResult(jobid, "out")
162
- ```
163
-
164
- **Note:** BLAST jobs are asynchronous. Check status before retrieving results.
165
-
166
- ### 5. Identifier Mapping
167
-
168
- Convert identifiers between different biological databases:
169
-
170
- ```python
171
- from bioservices import UniProt, KEGG
172
-
173
- # UniProt mapping (many database pairs supported)
174
- u = UniProt()
175
- results = u.mapping(
176
- fr="UniProtKB_AC-ID", # Source database
177
- to="KEGG", # Target database
178
- query="P43403" # Identifier(s) to convert
179
- )
180
-
181
- # KEGG gene ID → UniProt
182
- kegg_to_uniprot = u.mapping(fr="KEGG", to="UniProtKB_AC-ID", query="hsa:7535")
183
-
184
- # For compounds, use UniChem
185
- from bioservices import UniChem
186
- u = UniChem()
187
- chembl_from_kegg = u.get_compound_id_from_kegg("C11222")
188
- ```
189
-
190
- **Supported mappings (UniProt):**
191
- - UniProtKB ↔ KEGG
192
- - UniProtKB ↔ Ensembl
193
- - UniProtKB ↔ PDB
194
- - UniProtKB ↔ RefSeq
195
- - And many more (see `references/identifier_mapping.md`)
196
-
197
- ### 6. Gene Ontology Queries
198
-
199
- Access GO terms and annotations:
200
-
201
- ```python
202
- from bioservices import QuickGO
203
-
204
- g = QuickGO(verbose=False)
205
-
206
- # Retrieve GO term information
207
- term_info = g.Term("GO:0003824", frmt="obo")
208
-
209
- # Search annotations
210
- annotations = g.Annotation(protein="P43403", format="tsv")
211
- ```
212
-
213
- ### 7. Protein-Protein Interactions
214
-
215
- Query interaction databases via PSICQUIC. **PSICQUIC is not shipped by every
216
- release — it is absent from 1.16.0** — so import it defensively and fall back to
217
- `IntactComplex`, `OmniPath`, or `STRING` when it is missing:
218
-
219
- ```python
220
- from bioservices import PSICQUIC
221
-
222
- s = PSICQUIC(verbose=False)
223
-
224
- # Query specific database (e.g., MINT)
225
- interactions = s.query("mint", "ZAP70 AND species:9606")
226
-
227
- # List available interaction databases
228
- databases = s.activeDBs
229
- ```
230
-
231
- **Available databases:** MINT, IntAct, BioGRID, DIP, and 30+ others.
232
-
233
- ## Multi-Service Integration Workflows
234
-
235
- BioServices excels at combining multiple services for comprehensive analysis. Common integration patterns:
236
-
237
- ### Complete Protein Analysis Pipeline
238
-
239
- Execute a full protein characterization workflow:
240
-
241
- ```bash
242
- export NCBI_EMAIL=your.email@example.com
243
- python scripts/protein_analysis_workflow.py ZAP70_HUMAN
244
- # Or pass email as optional second argument if NCBI_EMAIL is unset
245
- python scripts/protein_analysis_workflow.py ZAP70_HUMAN your.email@example.com
246
- ```
247
-
248
- This script demonstrates:
249
- 1. UniProt search for protein entry
250
- 2. FASTA sequence retrieval
251
- 3. BLAST similarity search
252
- 4. KEGG pathway discovery
253
- 5. PSICQUIC interaction mapping
254
-
255
- ### Pathway Network Analysis
256
-
257
- Analyze all pathways for an organism:
258
-
259
- ```bash
260
- python scripts/pathway_analysis.py hsa output_directory/
261
- ```
262
-
263
- Extracts and analyzes:
264
- - All pathway IDs for organism
265
- - Protein-protein interactions per pathway
266
- - Interaction type distributions
267
- - Exports to CSV/SIF formats
268
-
269
- ### Cross-Database Compound Search
270
-
271
- Map compound identifiers across databases:
272
-
273
- ```bash
274
- python scripts/compound_cross_reference.py Geldanamycin
275
- ```
276
-
277
- Retrieves:
278
- - KEGG compound ID
279
- - ChEBI identifier
280
- - ChEMBL identifier
281
- - Basic compound properties
282
-
283
- ### Batch Identifier Conversion
284
-
285
- Convert multiple identifiers at once:
286
-
287
- ```bash
288
- python scripts/batch_id_converter.py input_ids.txt --from UniProtKB_AC-ID --to KEGG
289
- ```
290
-
291
- ## Best Practices
292
-
293
- ### Output Format Handling
294
-
295
- Different services return data in various formats:
296
- - **XML**: Parse using BeautifulSoup (most SOAP services)
297
- - **Tab-separated (TSV)**: Pandas DataFrames for tabular data
298
- - **Dictionary/JSON**: Direct Python manipulation
299
- - **FASTA**: BioPython integration for sequence analysis
300
-
301
- ### Rate Limiting and Verbosity
302
-
303
- Control API request behavior:
304
-
305
- ```python
306
- from bioservices import KEGG
307
-
308
- k = KEGG(verbose=False) # Suppress HTTP request details
309
- k.TIMEOUT = 30 # Adjust timeout for slow connections
310
- ```
311
-
312
- ### Error Handling
313
-
314
- Wrap service calls in try-except blocks:
315
-
316
- ```python
317
- try:
318
- results = u.search("ambiguous_query")
319
- if results:
320
- # Process results
321
- pass
322
- except Exception as e:
323
- print(f"Search failed: {e}")
324
- ```
325
-
326
- ### Organism Codes
327
-
328
- Use standard organism abbreviations:
329
- - `hsa`: Homo sapiens (human)
330
- - `mmu`: Mus musculus (mouse)
331
- - `dme`: Drosophila melanogaster
332
- - `sce`: Saccharomyces cerevisiae (yeast)
333
-
334
- List all organisms: `k.list("organism")` or `k.organismIds`
335
-
336
- ### Integration with Other Tools
337
-
338
- BioServices works well with:
339
- - **BioPython**: Sequence analysis on retrieved FASTA data
340
- - **Pandas**: Tabular data manipulation
341
- - **PyMOL**: 3D structure visualization (retrieve PDB IDs)
342
- - **NetworkX**: Network analysis of pathway interactions
343
- - **Galaxy**: Custom tool wrappers for workflow platforms
344
-
345
- ## Resources
346
-
347
- ### scripts/
348
-
349
- Executable Python scripts demonstrating complete workflows:
350
-
351
- - `protein_analysis_workflow.py`: End-to-end protein characterization
352
- - `pathway_analysis.py`: KEGG pathway discovery and network extraction
353
- - `compound_cross_reference.py`: Multi-database compound searching
354
- - `batch_id_converter.py`: Bulk identifier mapping utility
355
-
356
- Scripts can be executed directly or adapted for specific use cases.
357
-
358
- ### references/
359
-
360
- Detailed documentation loaded as needed:
361
-
362
- - `services_reference.md`: Comprehensive list of all 40+ services with methods
363
- - `workflow_patterns.md`: Detailed multi-step analysis workflows
364
- - `identifier_mapping.md`: Complete guide to cross-database ID conversion
365
-
366
- Load references when working with specific services or complex integration tasks.
367
-
368
- ## Installation
369
-
370
- ```bash
371
- uv pip install "bioservices==1.16.0"
372
- ```
373
-
374
- Dependencies are installed automatically. Upstream CI tests Python 3.9–3.12 ([PyPI](https://pypi.org/project/bioservices/), [docs](https://bioservices.readthedocs.io/)).
375
-
376
- ## Credentials
377
-
378
- Most services need no API key. Exceptions:
379
-
380
- | Service | Requirement |
381
- |---------|-------------|
382
- | NCBI BLAST | Contact email via `NCBI_EMAIL` or `email=` in `NCBIblast.run()` |
383
- | Some EBI services | Optional; check service docs if rate-limited |
384
-
385
- Set once per shell session:
386
-
387
- ```bash
388
- export NCBI_EMAIL=your.email@example.com
389
- ```
390
-
391
- Use a real institutional or lab address — NCBI may contact you about heavy BLAST usage.
392
-
393
- ## Additional Information
394
-
395
- For detailed API documentation and advanced features, refer to:
396
- - Official documentation: https://bioservices.readthedocs.io/
397
- - Source code: https://github.com/cokelaer/bioservices
398
- - Service-specific references in `references/services_reference.md`
399
-
@@ -1,198 +0,0 @@
1
- ---
2
- name: bulk-rnaseq
3
- description: End-to-end bulk RNA-seq orchestrator — takes raw FASTQ reads through QC and trimming (FastQC, fastp/Trim Galore), alignment and quantification (STAR, Salmon, featureCounts), assembles a gene-level counts matrix, then hands off to differential expression (pydeseq2), pathway/GSEA enrichment (pathway-enrichment), and publication figures (scientific-visualization). Use whenever the user has bulk RNA-seq reads or quant output and wants a complete, reproducible differential-expression workflow — e.g. "analyze my RNA-seq", "FASTQ to DESeq2", "run nf-core/rnaseq", "STAR/Salmon quantification", "build a counts matrix for DESeq2", or "go from reads to differentially expressed genes and enriched pathways". Routes between an nf-core/rnaseq (Nextflow) path and a standalone STAR/Salmon path, and covers experimental design, strandedness, and QC gates. For single-cell RNA-seq use the scanpy skill instead.
4
- license: MIT
5
- metadata:
6
- version: "1.0"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # Bulk RNA-seq
11
-
12
- ## Overview
13
-
14
- This skill orchestrates a complete, **defensible** bulk RNA-seq differential-expression study, from raw sequencing reads to enriched pathways and figures. It is a router, not a reimplementation: most stages already have dedicated skills in this repo, and this skill connects them in the right order, fills the one real gap (raw reads → a gene-level counts matrix), and enforces the design and QC decisions that determine whether the final result is trustworthy.
15
-
16
- "Defensible" means three things, applied throughout:
17
- - **Reproducible** — pinned pipeline/tool versions, containers where possible, recorded parameters, fixed random seeds.
18
- - **Quality-gated** — QC is inspected and acted on before, during, and after quantification, not skipped.
19
- - **Statistically sound** — adequate replication, a design that matches the biology, counts handled correctly, and FDR-controlled testing.
20
-
21
- The pipeline is: **FastQC/trim → align/quant (STAR/Salmon) → counts → DE (pydeseq2) → enrichment (pathway-enrichment) → figures**.
22
-
23
- ## When to Use This Skill
24
-
25
- Use this skill when the user wants to:
26
- - Go from FASTQ files (or a sequencing run) to differentially expressed genes and pathways.
27
- - Run or configure `nf-core/rnaseq`, or align/quantify with STAR, Salmon, or featureCounts.
28
- - Turn Salmon/STAR/featureCounts output into a counts matrix ready for DESeq2/PyDESeq2.
29
- - Design or sanity-check a bulk RNA-seq experiment (replicates, batch, strandedness) before committing compute.
30
- - Scope an end-to-end RNA-seq analysis and decide which tools and skills to chain.
31
-
32
- This is **bulk** RNA-seq (samples = biological specimens). For single-cell/nuclei data use `scanpy`; for the DE statistics alone use `pydeseq2`; for enrichment alone use `pathway-enrichment`.
33
-
34
- ## The Pipeline at a Glance
35
-
36
- ```mermaid
37
- flowchart TD
38
- fastq["Raw FASTQ + samplesheet"] --> qc["FastQC + MultiQC"]
39
- qc --> trim["Trim: fastp / Trim Galore"]
40
- trim --> align["Align + quant: STAR and/or Salmon"]
41
- align --> counts["Gene-level counts matrix"]
42
- counts --> de["Differential expression"]
43
- de --> enrich["Pathway / GSEA enrichment"]
44
- de --> fig["Figures"]
45
- enrich --> fig
46
- nfcore["nf-core/rnaseq via nextflow skill"] -.->|"path A"| align
47
- manual["Standalone recipes (this skill)"] -.->|"path B"| align
48
- bridge["build_counts_matrix.py (this skill)"] -.-> counts
49
- pydeseq2skill["pydeseq2 skill"] -.-> de
50
- pwskill["pathway-enrichment skill"] -.-> enrich
51
- vizskill["scientific-visualization skill"] -.-> fig
52
- ```
53
-
54
- ## Two Upstream Paths — Pick One
55
-
56
- The reads → counts stage can be run two ways. They produce equivalent gene counts; choose by context, then stay on that path.
57
-
58
- | Use **Path A — `nf-core/rnaseq`** when… | Use **Path B — standalone tools** when… |
59
- |------------------------------------------|------------------------------------------|
60
- | You want the field-standard, audited, citable pipeline with one command | You have a few samples and want to learn/inspect each step |
61
- | Many samples, or you'll scale to HPC/cloud | No Nextflow/containers available, or a constrained environment |
62
- | Reproducibility and a full MultiQC report matter most | You need a non-standard step the pipeline doesn't expose |
63
- | → Drive it through the **`nextflow`** skill | → Follow `references/upstream-manual.md` |
64
-
65
- When unsure, prefer **Path A**: `nf-core/rnaseq` already wires together FastQC → trimming → STAR/Salmon → quantification → tximport → MultiQC with sensible, reviewed defaults, which is the most defensible option. Path B exists for transparency and constrained setups.
66
-
67
- Both paths converge on a **gene-level counts matrix**, after which the workflow is identical.
68
-
69
- ## Setup
70
-
71
- ```bash
72
- # This skill's glue (bridge + handoffs) — Python
73
- uv pip install pytximport pandas
74
-
75
- # Downstream skills install their own deps:
76
- # pydeseq2 skill -> uv pip install pydeseq2
77
- # pathway-enrichment skill -> uv pip install gseapy gprofiler-official
78
-
79
- # Path A (nf-core): only Nextflow + a container engine are needed — see the `nextflow` skill.
80
-
81
- # Path B (standalone tools): install via bioconda. Pin versions for reproducibility.
82
- conda create -n rnaseq -c bioconda -c conda-forge \
83
- fastqc fastp trim-galore "star=2.7.11b" "salmon=1.10.3" subread multiqc
84
- ```
85
-
86
- Record the exact versions you use (pipeline revision, tool versions, reference genome + annotation release) — they belong in the methods section and make the analysis reproducible.
87
-
88
- ## Quick Start
89
-
90
- ### Path A — nf-core/rnaseq (recommended)
91
-
92
- ```bash
93
- # 0. Validate the samplesheet first (catches the most common failures early)
94
- python scripts/validate_samplesheet.py --samplesheet samplesheet.csv
95
-
96
- # 1. Smoke-test the environment with tiny bundled data
97
- nextflow run nf-core/rnaseq -r 3.26.0 -profile test,docker --outdir test_results
98
-
99
- # 2. Real run: pin the revision, pick an aligner, pass a samplesheet + reference
100
- nextflow run nf-core/rnaseq -r 3.26.0 \
101
- -profile docker \
102
- --input samplesheet.csv \
103
- --genome GRCh38 \
104
- --aligner star_salmon \
105
- --outdir results \
106
- -resume
107
- ```
108
-
109
- `nf-core/rnaseq` runs tximport internally, so gene counts come out **already merged** — no bridge script needed. Use `results/star_salmon/salmon.merged.gene_counts_length_scaled.tsv` for DE. Samplesheet format, aligner choice, and outputs: `references/upstream-nfcore.md`. For engine/HPC/cloud/container detail, use the **`nextflow`** skill.
110
-
111
- ### Path B — standalone STAR/Salmon (abbreviated)
112
-
113
- ```bash
114
- fastqc -o qc/ reads/*.fastq.gz # 1. QC raw reads
115
- fastp -i s1_R1.fq.gz -I s1_R2.fq.gz \
116
- -o s1_R1.trim.fq.gz -O s1_R2.trim.fq.gz \
117
- --thread 4 -j s1.fastp.json # 2. Trim adapters/low-quality
118
- salmon quant -i salmon_index -l A \
119
- -1 s1_R1.trim.fq.gz -2 s1_R2.trim.fq.gz \
120
- --gcBias --seqBias -p 8 -o quant/s1 # 3. Quantify (per sample)
121
- ```
122
-
123
- Full recipes (FastQC, fastp/Trim Galore, STAR index+align+`--quantMode GeneCounts`, Salmon decoy-aware index, featureCounts, strandedness): `references/upstream-manual.md`.
124
-
125
- ### Counts → DE → enrichment (both paths)
126
-
127
- ```bash
128
- # Path B only: assemble a gene x sample counts matrix + metadata template for PyDESeq2
129
- python scripts/build_counts_matrix.py --from salmon \
130
- --quant-dir quant/ --tx2gene tx2gene.tsv --output-dir counts/
131
-
132
- # Then hand off (see the dedicated skills):
133
- # pydeseq2: counts.csv + metadata.csv -> DE table (log2FC, padj, stat)
134
- # pathway-enrichment: rank by `stat` (GSEA) or padj+|LFC| hit list (ORA)
135
- # scientific-visualization / matplotlib: volcano, MA, heatmap, PCA, enrichment dotplot
136
- ```
137
-
138
- ## Stage-by-Stage Workflow
139
-
140
- Work top to bottom. Each stage names the skill or file that owns the detail. Don't skip the design/QC stages — they are where bulk RNA-seq studies most often go wrong.
141
-
142
- 1. **Design & sample sheet.** Confirm ≥3 biological replicates per group, identify batch/confounders, and choose the comparison(s). Build the samplesheet and validate it with `scripts/validate_samplesheet.py`. Rationale and rules: `references/design-and-qc.md`.
143
- 2. **Raw-read QC.** FastQC per file; aggregate with MultiQC. Check per-base quality, adapter content, duplication, and over-representation. Thresholds: `references/design-and-qc.md`.
144
- 3. **Trimming.** Remove adapters and low-quality tails (via `fastp` or `Trim Galore`). Re-run FastQC to confirm. Recipes: `references/upstream-manual.md` (Path A does this for you).
145
- 4. **Align / quantify.** STAR (genome alignment + `--quantMode GeneCounts`) and/or Salmon (transcript quasi-mapping, decoy-aware). Determine strandedness — it is easy to get wrong and silently halves your counts. Detail: `references/upstream-manual.md`; pipeline params: `references/upstream-nfcore.md`.
146
- 5. **Build the counts matrix.** Turn quant output into a gene × sample integer matrix and a metadata template (`scripts/build_counts_matrix.py`). The estimated-count and gene-ID-mapping nuances live in `references/counts-and-handoff.md`.
147
- 6. **Differential expression → `pydeseq2` skill.** Load `counts.csv` + `metadata.csv`, set the design (e.g. `~batch + condition`), fit, and test with FDR control. Inspect the PCA and p-value histogram as QC.
148
- 7. **Enrichment → `pathway-enrichment` skill.** For GSEA, rank the *full* gene list by the DESeq2 `stat`; for ORA, pass the thresholded hit list (padj < 0.05, optionally |log2FC| > 1). Map gene IDs to symbols first.
149
- 8. **Figures → `scientific-visualization` skill.** Volcano, MA, sample-distance heatmap, PCA, and enrichment dotplots, plus the MultiQC report for the QC narrative.
150
-
151
- ## The counts → DE bridge (the key glue)
152
-
153
- This is the one stage with no upstream/downstream skill, so this skill owns it. `scripts/build_counts_matrix.py` converts quant output into exactly what `pydeseq2` expects:
154
-
155
- - **Salmon** (`--from salmon`): aggregates per-sample `quant.sf` to gene level with `pytximport` using `counts_from_abundance="length_scaled_tpm"` (the right choice for gene-level DE), needs a `tx2gene` map.
156
- - **STAR** (`--from star`): reads each `ReadsPerGene.out.tab`, selecting the column for your `--strandedness` (unstranded/forward/reverse).
157
- - **featureCounts** (`--from featurecounts`): parses the combined `featureCounts` matrix.
158
-
159
- It writes `counts.csv` (genes × samples, integers) and `metadata_template.csv` (one row per sample) for you to fill in. **Salmon/RSEM counts are estimates (non-integer); they are rounded to integers** because PyDESeq2 requires integer counts — see `references/counts-and-handoff.md` for why this is acceptable with `length_scaled_tpm` and how it differs from the offset-based DESeq2+tximport route. That reference also covers Ensembl→symbol mapping (needed before enrichment) and the exact orientation PyDESeq2 wants.
160
-
161
- ## Common Pitfalls
162
-
163
- These cause most wrong or irreproducible bulk RNA-seq results:
164
-
165
- 1. **Too few replicates.** <3 biological replicates per group gives almost no power and unstable dispersion estimates. More replicates beat deeper sequencing.
166
- 2. **Confounded batch and condition.** If every treated sample was processed on a different day/lane than controls, the effect is unrecoverable. Randomize, and model known batches (`~batch + condition`). See `references/design-and-qc.md`.
167
- 3. **Wrong strandedness.** Choosing the wrong STAR column or featureCounts `-s`/Salmon library type silently discards ~half the reads. Use Salmon `-l A` or infer strandedness, and verify the assigned-reads fraction.
168
- 4. **Feeding TPM/FPKM to DESeq2.** DESeq2 needs raw (or length-scaled) **counts**, never TPM/FPKM/normalized values. The bridge handles this.
169
- 5. **Non-integer counts.** PyDESeq2 requires integers; round Salmon estimates (the bridge does this).
170
- 6. **Gene-ID mismatch into enrichment.** DESeq2 output is often Ensembl IDs; Enrichr/MSigDB want symbols. Map IDs before `pathway-enrichment` or "nothing is significant".
171
- 7. **Skipping post-quant QC.** Always look at the PCA and sample-distance heatmap before trusting DE — they expose swapped labels, outliers, and hidden batches.
172
- 8. **Mixing aligners across samples.** Quantify every sample with the same tool, version, reference, and parameters.
173
- 9. **Unpinned versions.** "latest" pipelines/genomes make results unreproducible; pin `-r`, tool versions, and the genome/annotation release.
174
-
175
- ## Integration with Other Skills
176
-
177
- - **Upstream execution:** `nextflow` (runs `nf-core/rnaseq`, Path A; HPC/cloud/containers).
178
- - **Reference data / gene IDs:** `gget` (`gget ref` for genome+GTF, `gget info`/`gget search` for ID mapping), `database-lookup` (Ensembl/NCBI), `biopython`/`pysam` (FASTA/BAM handling).
179
- - **Differential expression:** `pydeseq2` (the DE engine this skill hands counts to).
180
- - **Enrichment:** `pathway-enrichment` (ORA + GSEA; its `scripts/run_enrichment.py` reads a DESeq2 results CSV directly).
181
- - **Figures & reporting:** `scientific-visualization`, `matplotlib`, `seaborn`; `scientific-writing` for the methods/results narrative.
182
- - **Related but distinct:** `scanpy` (single-cell), `statistical-analysis` (multiple-testing depth).
183
-
184
- ## Reference Files
185
-
186
- Read the relevant file when you need depth — each is self-contained:
187
-
188
- - `references/upstream-nfcore.md` — Path A: samplesheet format, `--aligner`/`--pseudo_aligner` choice, key params, the `salmon.merged.gene_counts*.tsv` outputs, MultiQC, and what to hand to `pydeseq2`.
189
- - `references/upstream-manual.md` — Path B: FastQC, fastp/Trim Galore, STAR genome index + alignment + `--quantMode GeneCounts`, Salmon decoy-aware index + `quant`, featureCounts, and how to determine strandedness.
190
- - `references/counts-and-handoff.md` — turning quant output into PyDESeq2-ready `counts.csv`/`metadata.csv` (pytximport, STAR column selection, featureCounts), the integer/estimated-count nuance, Ensembl→symbol mapping, and the DE→enrichment rank/hit-list recipe.
191
- - `references/design-and-qc.md` — experimental design (replication, batch, confounding, design formulas) and QC-metric interpretation (mapping rate, duplication, rRNA, complexity, PCA/outliers) — the defensible-pipeline backbone.
192
-
193
- ## Resources
194
-
195
- - nf-core/rnaseq: https://nf-co.re/rnaseq · STAR: https://github.com/alexdobin/STAR · Salmon: https://salmon.readthedocs.io
196
- - fastp: https://github.com/OpenGene/fastp · Trim Galore: https://github.com/FelixKrueger/TrimGalore · MultiQC: https://multiqc.info
197
- - pytximport: https://pytximport.complextissue.com · featureCounts (Subread): https://subread.sourceforge.net
198
- - Method background: Love et al. 2014 (DESeq2) DOI 10.1186/s13059-014-0550-8 · Soneson et al. 2015 (tximport) DOI 10.12688/f1000research.7563.2