@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,399 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: bioservices
|
|
3
|
-
description: Unified Python interface to 40+ bioinformatics services. Use when querying multiple databases (UniProt, KEGG, ChEMBL, Reactome) in a single workflow with consistent API. Best for cross-database analysis, ID mapping across services. For quick single-database lookups use gget; for sequence/file manipulation use biopython.
|
|
4
|
-
license: GPLv3 license
|
|
5
|
-
allowed-tools: Read Write Edit Bash
|
|
6
|
-
compatibility: Requires Python 3.9–3.12 and internet access to 40+ bioinformatics web APIs. NCBI BLAST requires a contact email (`NCBI_EMAIL` env var or explicit parameter).
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.3"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
openclaw:
|
|
11
|
-
envVars:
|
|
12
|
-
- name: NCBI_EMAIL
|
|
13
|
-
required: false
|
|
14
|
-
description: Email for NCBI service identification.
|
|
15
|
-
---
|
|
16
|
-
|
|
17
|
-
# BioServices
|
|
18
|
-
|
|
19
|
-
## Overview
|
|
20
|
-
|
|
21
|
-
BioServices is a Python package providing programmatic access to approximately 40 bioinformatics web services and databases. Retrieve biological data, perform cross-database queries, map identifiers, analyze sequences, and integrate multiple biological resources in Python workflows. The package handles both REST and SOAP/WSDL protocols transparently.
|
|
22
|
-
|
|
23
|
-
**Version note:** Examples target **bioservices 1.16.0** (PyPI, Mar 2026). Requires **Python 3.9–3.12**. UniProt REST changes in mid-2022 (bioservices ≥1.10) mainly affect tabular `columns` names — see upstream `_legacy_names` if parsing breaks. ChEMBL wrappers changed at 1.6.0 (2018 API); use `get_similarity`, `get_substructure`, `get_molecule` instead of pre-1.6 method names.
|
|
24
|
-
|
|
25
|
-
## When to Use This Skill
|
|
26
|
-
|
|
27
|
-
This skill should be used when:
|
|
28
|
-
- Retrieving protein sequences, annotations, or structures from UniProt, PDB, Pfam
|
|
29
|
-
- Analyzing metabolic pathways and gene functions via KEGG or Reactome
|
|
30
|
-
- Searching compound databases (ChEBI, ChEMBL, PubChem) for chemical information
|
|
31
|
-
- Converting identifiers between different biological databases (KEGG↔UniProt, compound IDs)
|
|
32
|
-
- Running sequence similarity searches (BLAST, MUSCLE alignment)
|
|
33
|
-
- Querying gene ontology terms (QuickGO, GO annotations)
|
|
34
|
-
- Accessing protein-protein interaction data (PSICQUIC, IntactComplex)
|
|
35
|
-
- Mining genomic data (BioMart, ArrayExpress, ENA)
|
|
36
|
-
- Integrating data from multiple bioinformatics resources in a single workflow
|
|
37
|
-
|
|
38
|
-
## Core Capabilities
|
|
39
|
-
|
|
40
|
-
### 1. Protein Analysis
|
|
41
|
-
|
|
42
|
-
Retrieve protein information, sequences, and functional annotations:
|
|
43
|
-
|
|
44
|
-
```python
|
|
45
|
-
from bioservices import UniProt
|
|
46
|
-
|
|
47
|
-
u = UniProt(verbose=False)
|
|
48
|
-
|
|
49
|
-
# Search for protein by name
|
|
50
|
-
results = u.search("ZAP70_HUMAN", frmt="tab", columns="id,genes,organism")
|
|
51
|
-
|
|
52
|
-
# Retrieve FASTA sequence
|
|
53
|
-
sequence = u.retrieve("P43403", "fasta")
|
|
54
|
-
|
|
55
|
-
# Map identifiers between databases
|
|
56
|
-
kegg_ids = u.mapping(fr="UniProtKB_AC-ID", to="KEGG", query="P43403")
|
|
57
|
-
```
|
|
58
|
-
|
|
59
|
-
**Key methods:**
|
|
60
|
-
- `search()`: Query UniProt with flexible search terms
|
|
61
|
-
- `retrieve()`: Get protein entries in various formats (FASTA, XML, tab)
|
|
62
|
-
- `mapping()`: Convert identifiers between databases
|
|
63
|
-
|
|
64
|
-
Reference: `references/services_reference.md` for complete UniProt API details.
|
|
65
|
-
|
|
66
|
-
### 2. Pathway Discovery and Analysis
|
|
67
|
-
|
|
68
|
-
Access KEGG pathway information for genes and organisms:
|
|
69
|
-
|
|
70
|
-
```python
|
|
71
|
-
from bioservices import KEGG
|
|
72
|
-
|
|
73
|
-
k = KEGG()
|
|
74
|
-
k.organism = "hsa" # Set to human
|
|
75
|
-
|
|
76
|
-
# Search for organisms
|
|
77
|
-
k.lookfor_organism("droso") # Find Drosophila species
|
|
78
|
-
|
|
79
|
-
# Find pathways by name
|
|
80
|
-
k.lookfor_pathway("B cell") # Returns matching pathway IDs
|
|
81
|
-
|
|
82
|
-
# Get pathways containing specific genes
|
|
83
|
-
pathways = k.get_pathway_by_gene("7535", "hsa") # ZAP70 gene
|
|
84
|
-
|
|
85
|
-
# Retrieve and parse pathway data
|
|
86
|
-
data = k.get("hsa04660")
|
|
87
|
-
parsed = k.parse(data)
|
|
88
|
-
|
|
89
|
-
# Extract pathway interactions
|
|
90
|
-
interactions = k.parse_kgml_pathway("hsa04660")
|
|
91
|
-
relations = interactions['relations'] # Protein-protein interactions
|
|
92
|
-
|
|
93
|
-
# Convert to Simple Interaction Format
|
|
94
|
-
sif_data = k.pathway2sif("hsa04660")
|
|
95
|
-
```
|
|
96
|
-
|
|
97
|
-
**Key methods:**
|
|
98
|
-
- `lookfor_organism()`, `lookfor_pathway()`: Search by name
|
|
99
|
-
- `get_pathway_by_gene()`: Find pathways containing genes
|
|
100
|
-
- `parse_kgml_pathway()`: Extract structured pathway data
|
|
101
|
-
- `pathway2sif()`: Get protein interaction networks
|
|
102
|
-
|
|
103
|
-
Reference: `references/workflow_patterns.md` for complete pathway analysis workflows.
|
|
104
|
-
|
|
105
|
-
### 3. Compound Database Searches
|
|
106
|
-
|
|
107
|
-
Search and cross-reference compounds across multiple databases:
|
|
108
|
-
|
|
109
|
-
```python
|
|
110
|
-
from bioservices import KEGG, UniChem
|
|
111
|
-
|
|
112
|
-
k = KEGG()
|
|
113
|
-
|
|
114
|
-
# Search compounds by name
|
|
115
|
-
results = k.find("compound", "Geldanamycin") # Returns cpd:C11222
|
|
116
|
-
|
|
117
|
-
# Get compound information with database links
|
|
118
|
-
compound_info = k.get("cpd:C11222") # Includes ChEBI links
|
|
119
|
-
|
|
120
|
-
# Cross-reference KEGG → ChEMBL using UniChem
|
|
121
|
-
u = UniChem()
|
|
122
|
-
chembl_id = u.get_compound_id_from_kegg("C11222") # Returns CHEMBL278315
|
|
123
|
-
```
|
|
124
|
-
|
|
125
|
-
**Version caveat:** the per-source `get_compound_id_from_*` helpers are gone from
|
|
126
|
-
bioservices 1.16.0 — check `hasattr(u, "get_compound_id_from_kegg")` first, and
|
|
127
|
-
otherwise use the current UniChem API (`u.get_compounds(compound, source_type)`
|
|
128
|
-
and read `res["compounds"][0]["sources"]`). ChEMBL lookups follow the same rule:
|
|
129
|
-
`get_molecule`, not the pre-1.6 `get_compound_by_chemblId`.
|
|
130
|
-
|
|
131
|
-
**Common workflow:**
|
|
132
|
-
1. Search compound by name in KEGG
|
|
133
|
-
2. Extract KEGG compound ID
|
|
134
|
-
3. Use UniChem for KEGG → ChEMBL mapping
|
|
135
|
-
4. ChEBI IDs are often provided in KEGG entries
|
|
136
|
-
|
|
137
|
-
Reference: `references/identifier_mapping.md` for complete cross-database mapping guide.
|
|
138
|
-
|
|
139
|
-
### 4. Sequence Analysis
|
|
140
|
-
|
|
141
|
-
Run BLAST searches and sequence alignments. NCBI requires a contact email — prefer the `NCBI_EMAIL` environment variable (same convention as BioPython Entrez and other repo skills):
|
|
142
|
-
|
|
143
|
-
```python
|
|
144
|
-
import os
|
|
145
|
-
from bioservices import NCBIblast
|
|
146
|
-
|
|
147
|
-
s = NCBIblast(verbose=False)
|
|
148
|
-
email = os.environ["NCBI_EMAIL"] # set before running: export NCBI_EMAIL=you@lab.org
|
|
149
|
-
|
|
150
|
-
# Run BLASTP against UniProtKB
|
|
151
|
-
jobid = s.run(
|
|
152
|
-
program="blastp",
|
|
153
|
-
sequence=protein_sequence,
|
|
154
|
-
stype="protein",
|
|
155
|
-
database="uniprotkb",
|
|
156
|
-
email=email,
|
|
157
|
-
)
|
|
158
|
-
|
|
159
|
-
# Check job status and retrieve results
|
|
160
|
-
s.getStatus(jobid)
|
|
161
|
-
results = s.getResult(jobid, "out")
|
|
162
|
-
```
|
|
163
|
-
|
|
164
|
-
**Note:** BLAST jobs are asynchronous. Check status before retrieving results.
|
|
165
|
-
|
|
166
|
-
### 5. Identifier Mapping
|
|
167
|
-
|
|
168
|
-
Convert identifiers between different biological databases:
|
|
169
|
-
|
|
170
|
-
```python
|
|
171
|
-
from bioservices import UniProt, KEGG
|
|
172
|
-
|
|
173
|
-
# UniProt mapping (many database pairs supported)
|
|
174
|
-
u = UniProt()
|
|
175
|
-
results = u.mapping(
|
|
176
|
-
fr="UniProtKB_AC-ID", # Source database
|
|
177
|
-
to="KEGG", # Target database
|
|
178
|
-
query="P43403" # Identifier(s) to convert
|
|
179
|
-
)
|
|
180
|
-
|
|
181
|
-
# KEGG gene ID → UniProt
|
|
182
|
-
kegg_to_uniprot = u.mapping(fr="KEGG", to="UniProtKB_AC-ID", query="hsa:7535")
|
|
183
|
-
|
|
184
|
-
# For compounds, use UniChem
|
|
185
|
-
from bioservices import UniChem
|
|
186
|
-
u = UniChem()
|
|
187
|
-
chembl_from_kegg = u.get_compound_id_from_kegg("C11222")
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
**Supported mappings (UniProt):**
|
|
191
|
-
- UniProtKB ↔ KEGG
|
|
192
|
-
- UniProtKB ↔ Ensembl
|
|
193
|
-
- UniProtKB ↔ PDB
|
|
194
|
-
- UniProtKB ↔ RefSeq
|
|
195
|
-
- And many more (see `references/identifier_mapping.md`)
|
|
196
|
-
|
|
197
|
-
### 6. Gene Ontology Queries
|
|
198
|
-
|
|
199
|
-
Access GO terms and annotations:
|
|
200
|
-
|
|
201
|
-
```python
|
|
202
|
-
from bioservices import QuickGO
|
|
203
|
-
|
|
204
|
-
g = QuickGO(verbose=False)
|
|
205
|
-
|
|
206
|
-
# Retrieve GO term information
|
|
207
|
-
term_info = g.Term("GO:0003824", frmt="obo")
|
|
208
|
-
|
|
209
|
-
# Search annotations
|
|
210
|
-
annotations = g.Annotation(protein="P43403", format="tsv")
|
|
211
|
-
```
|
|
212
|
-
|
|
213
|
-
### 7. Protein-Protein Interactions
|
|
214
|
-
|
|
215
|
-
Query interaction databases via PSICQUIC. **PSICQUIC is not shipped by every
|
|
216
|
-
release — it is absent from 1.16.0** — so import it defensively and fall back to
|
|
217
|
-
`IntactComplex`, `OmniPath`, or `STRING` when it is missing:
|
|
218
|
-
|
|
219
|
-
```python
|
|
220
|
-
from bioservices import PSICQUIC
|
|
221
|
-
|
|
222
|
-
s = PSICQUIC(verbose=False)
|
|
223
|
-
|
|
224
|
-
# Query specific database (e.g., MINT)
|
|
225
|
-
interactions = s.query("mint", "ZAP70 AND species:9606")
|
|
226
|
-
|
|
227
|
-
# List available interaction databases
|
|
228
|
-
databases = s.activeDBs
|
|
229
|
-
```
|
|
230
|
-
|
|
231
|
-
**Available databases:** MINT, IntAct, BioGRID, DIP, and 30+ others.
|
|
232
|
-
|
|
233
|
-
## Multi-Service Integration Workflows
|
|
234
|
-
|
|
235
|
-
BioServices excels at combining multiple services for comprehensive analysis. Common integration patterns:
|
|
236
|
-
|
|
237
|
-
### Complete Protein Analysis Pipeline
|
|
238
|
-
|
|
239
|
-
Execute a full protein characterization workflow:
|
|
240
|
-
|
|
241
|
-
```bash
|
|
242
|
-
export NCBI_EMAIL=your.email@example.com
|
|
243
|
-
python scripts/protein_analysis_workflow.py ZAP70_HUMAN
|
|
244
|
-
# Or pass email as optional second argument if NCBI_EMAIL is unset
|
|
245
|
-
python scripts/protein_analysis_workflow.py ZAP70_HUMAN your.email@example.com
|
|
246
|
-
```
|
|
247
|
-
|
|
248
|
-
This script demonstrates:
|
|
249
|
-
1. UniProt search for protein entry
|
|
250
|
-
2. FASTA sequence retrieval
|
|
251
|
-
3. BLAST similarity search
|
|
252
|
-
4. KEGG pathway discovery
|
|
253
|
-
5. PSICQUIC interaction mapping
|
|
254
|
-
|
|
255
|
-
### Pathway Network Analysis
|
|
256
|
-
|
|
257
|
-
Analyze all pathways for an organism:
|
|
258
|
-
|
|
259
|
-
```bash
|
|
260
|
-
python scripts/pathway_analysis.py hsa output_directory/
|
|
261
|
-
```
|
|
262
|
-
|
|
263
|
-
Extracts and analyzes:
|
|
264
|
-
- All pathway IDs for organism
|
|
265
|
-
- Protein-protein interactions per pathway
|
|
266
|
-
- Interaction type distributions
|
|
267
|
-
- Exports to CSV/SIF formats
|
|
268
|
-
|
|
269
|
-
### Cross-Database Compound Search
|
|
270
|
-
|
|
271
|
-
Map compound identifiers across databases:
|
|
272
|
-
|
|
273
|
-
```bash
|
|
274
|
-
python scripts/compound_cross_reference.py Geldanamycin
|
|
275
|
-
```
|
|
276
|
-
|
|
277
|
-
Retrieves:
|
|
278
|
-
- KEGG compound ID
|
|
279
|
-
- ChEBI identifier
|
|
280
|
-
- ChEMBL identifier
|
|
281
|
-
- Basic compound properties
|
|
282
|
-
|
|
283
|
-
### Batch Identifier Conversion
|
|
284
|
-
|
|
285
|
-
Convert multiple identifiers at once:
|
|
286
|
-
|
|
287
|
-
```bash
|
|
288
|
-
python scripts/batch_id_converter.py input_ids.txt --from UniProtKB_AC-ID --to KEGG
|
|
289
|
-
```
|
|
290
|
-
|
|
291
|
-
## Best Practices
|
|
292
|
-
|
|
293
|
-
### Output Format Handling
|
|
294
|
-
|
|
295
|
-
Different services return data in various formats:
|
|
296
|
-
- **XML**: Parse using BeautifulSoup (most SOAP services)
|
|
297
|
-
- **Tab-separated (TSV)**: Pandas DataFrames for tabular data
|
|
298
|
-
- **Dictionary/JSON**: Direct Python manipulation
|
|
299
|
-
- **FASTA**: BioPython integration for sequence analysis
|
|
300
|
-
|
|
301
|
-
### Rate Limiting and Verbosity
|
|
302
|
-
|
|
303
|
-
Control API request behavior:
|
|
304
|
-
|
|
305
|
-
```python
|
|
306
|
-
from bioservices import KEGG
|
|
307
|
-
|
|
308
|
-
k = KEGG(verbose=False) # Suppress HTTP request details
|
|
309
|
-
k.TIMEOUT = 30 # Adjust timeout for slow connections
|
|
310
|
-
```
|
|
311
|
-
|
|
312
|
-
### Error Handling
|
|
313
|
-
|
|
314
|
-
Wrap service calls in try-except blocks:
|
|
315
|
-
|
|
316
|
-
```python
|
|
317
|
-
try:
|
|
318
|
-
results = u.search("ambiguous_query")
|
|
319
|
-
if results:
|
|
320
|
-
# Process results
|
|
321
|
-
pass
|
|
322
|
-
except Exception as e:
|
|
323
|
-
print(f"Search failed: {e}")
|
|
324
|
-
```
|
|
325
|
-
|
|
326
|
-
### Organism Codes
|
|
327
|
-
|
|
328
|
-
Use standard organism abbreviations:
|
|
329
|
-
- `hsa`: Homo sapiens (human)
|
|
330
|
-
- `mmu`: Mus musculus (mouse)
|
|
331
|
-
- `dme`: Drosophila melanogaster
|
|
332
|
-
- `sce`: Saccharomyces cerevisiae (yeast)
|
|
333
|
-
|
|
334
|
-
List all organisms: `k.list("organism")` or `k.organismIds`
|
|
335
|
-
|
|
336
|
-
### Integration with Other Tools
|
|
337
|
-
|
|
338
|
-
BioServices works well with:
|
|
339
|
-
- **BioPython**: Sequence analysis on retrieved FASTA data
|
|
340
|
-
- **Pandas**: Tabular data manipulation
|
|
341
|
-
- **PyMOL**: 3D structure visualization (retrieve PDB IDs)
|
|
342
|
-
- **NetworkX**: Network analysis of pathway interactions
|
|
343
|
-
- **Galaxy**: Custom tool wrappers for workflow platforms
|
|
344
|
-
|
|
345
|
-
## Resources
|
|
346
|
-
|
|
347
|
-
### scripts/
|
|
348
|
-
|
|
349
|
-
Executable Python scripts demonstrating complete workflows:
|
|
350
|
-
|
|
351
|
-
- `protein_analysis_workflow.py`: End-to-end protein characterization
|
|
352
|
-
- `pathway_analysis.py`: KEGG pathway discovery and network extraction
|
|
353
|
-
- `compound_cross_reference.py`: Multi-database compound searching
|
|
354
|
-
- `batch_id_converter.py`: Bulk identifier mapping utility
|
|
355
|
-
|
|
356
|
-
Scripts can be executed directly or adapted for specific use cases.
|
|
357
|
-
|
|
358
|
-
### references/
|
|
359
|
-
|
|
360
|
-
Detailed documentation loaded as needed:
|
|
361
|
-
|
|
362
|
-
- `services_reference.md`: Comprehensive list of all 40+ services with methods
|
|
363
|
-
- `workflow_patterns.md`: Detailed multi-step analysis workflows
|
|
364
|
-
- `identifier_mapping.md`: Complete guide to cross-database ID conversion
|
|
365
|
-
|
|
366
|
-
Load references when working with specific services or complex integration tasks.
|
|
367
|
-
|
|
368
|
-
## Installation
|
|
369
|
-
|
|
370
|
-
```bash
|
|
371
|
-
uv pip install "bioservices==1.16.0"
|
|
372
|
-
```
|
|
373
|
-
|
|
374
|
-
Dependencies are installed automatically. Upstream CI tests Python 3.9–3.12 ([PyPI](https://pypi.org/project/bioservices/), [docs](https://bioservices.readthedocs.io/)).
|
|
375
|
-
|
|
376
|
-
## Credentials
|
|
377
|
-
|
|
378
|
-
Most services need no API key. Exceptions:
|
|
379
|
-
|
|
380
|
-
| Service | Requirement |
|
|
381
|
-
|---------|-------------|
|
|
382
|
-
| NCBI BLAST | Contact email via `NCBI_EMAIL` or `email=` in `NCBIblast.run()` |
|
|
383
|
-
| Some EBI services | Optional; check service docs if rate-limited |
|
|
384
|
-
|
|
385
|
-
Set once per shell session:
|
|
386
|
-
|
|
387
|
-
```bash
|
|
388
|
-
export NCBI_EMAIL=your.email@example.com
|
|
389
|
-
```
|
|
390
|
-
|
|
391
|
-
Use a real institutional or lab address — NCBI may contact you about heavy BLAST usage.
|
|
392
|
-
|
|
393
|
-
## Additional Information
|
|
394
|
-
|
|
395
|
-
For detailed API documentation and advanced features, refer to:
|
|
396
|
-
- Official documentation: https://bioservices.readthedocs.io/
|
|
397
|
-
- Source code: https://github.com/cokelaer/bioservices
|
|
398
|
-
- Service-specific references in `references/services_reference.md`
|
|
399
|
-
|
|
@@ -1,198 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: bulk-rnaseq
|
|
3
|
-
description: End-to-end bulk RNA-seq orchestrator — takes raw FASTQ reads through QC and trimming (FastQC, fastp/Trim Galore), alignment and quantification (STAR, Salmon, featureCounts), assembles a gene-level counts matrix, then hands off to differential expression (pydeseq2), pathway/GSEA enrichment (pathway-enrichment), and publication figures (scientific-visualization). Use whenever the user has bulk RNA-seq reads or quant output and wants a complete, reproducible differential-expression workflow — e.g. "analyze my RNA-seq", "FASTQ to DESeq2", "run nf-core/rnaseq", "STAR/Salmon quantification", "build a counts matrix for DESeq2", or "go from reads to differentially expressed genes and enriched pathways". Routes between an nf-core/rnaseq (Nextflow) path and a standalone STAR/Salmon path, and covers experimental design, strandedness, and QC gates. For single-cell RNA-seq use the scanpy skill instead.
|
|
4
|
-
license: MIT
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.0"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Bulk RNA-seq
|
|
11
|
-
|
|
12
|
-
## Overview
|
|
13
|
-
|
|
14
|
-
This skill orchestrates a complete, **defensible** bulk RNA-seq differential-expression study, from raw sequencing reads to enriched pathways and figures. It is a router, not a reimplementation: most stages already have dedicated skills in this repo, and this skill connects them in the right order, fills the one real gap (raw reads → a gene-level counts matrix), and enforces the design and QC decisions that determine whether the final result is trustworthy.
|
|
15
|
-
|
|
16
|
-
"Defensible" means three things, applied throughout:
|
|
17
|
-
- **Reproducible** — pinned pipeline/tool versions, containers where possible, recorded parameters, fixed random seeds.
|
|
18
|
-
- **Quality-gated** — QC is inspected and acted on before, during, and after quantification, not skipped.
|
|
19
|
-
- **Statistically sound** — adequate replication, a design that matches the biology, counts handled correctly, and FDR-controlled testing.
|
|
20
|
-
|
|
21
|
-
The pipeline is: **FastQC/trim → align/quant (STAR/Salmon) → counts → DE (pydeseq2) → enrichment (pathway-enrichment) → figures**.
|
|
22
|
-
|
|
23
|
-
## When to Use This Skill
|
|
24
|
-
|
|
25
|
-
Use this skill when the user wants to:
|
|
26
|
-
- Go from FASTQ files (or a sequencing run) to differentially expressed genes and pathways.
|
|
27
|
-
- Run or configure `nf-core/rnaseq`, or align/quantify with STAR, Salmon, or featureCounts.
|
|
28
|
-
- Turn Salmon/STAR/featureCounts output into a counts matrix ready for DESeq2/PyDESeq2.
|
|
29
|
-
- Design or sanity-check a bulk RNA-seq experiment (replicates, batch, strandedness) before committing compute.
|
|
30
|
-
- Scope an end-to-end RNA-seq analysis and decide which tools and skills to chain.
|
|
31
|
-
|
|
32
|
-
This is **bulk** RNA-seq (samples = biological specimens). For single-cell/nuclei data use `scanpy`; for the DE statistics alone use `pydeseq2`; for enrichment alone use `pathway-enrichment`.
|
|
33
|
-
|
|
34
|
-
## The Pipeline at a Glance
|
|
35
|
-
|
|
36
|
-
```mermaid
|
|
37
|
-
flowchart TD
|
|
38
|
-
fastq["Raw FASTQ + samplesheet"] --> qc["FastQC + MultiQC"]
|
|
39
|
-
qc --> trim["Trim: fastp / Trim Galore"]
|
|
40
|
-
trim --> align["Align + quant: STAR and/or Salmon"]
|
|
41
|
-
align --> counts["Gene-level counts matrix"]
|
|
42
|
-
counts --> de["Differential expression"]
|
|
43
|
-
de --> enrich["Pathway / GSEA enrichment"]
|
|
44
|
-
de --> fig["Figures"]
|
|
45
|
-
enrich --> fig
|
|
46
|
-
nfcore["nf-core/rnaseq via nextflow skill"] -.->|"path A"| align
|
|
47
|
-
manual["Standalone recipes (this skill)"] -.->|"path B"| align
|
|
48
|
-
bridge["build_counts_matrix.py (this skill)"] -.-> counts
|
|
49
|
-
pydeseq2skill["pydeseq2 skill"] -.-> de
|
|
50
|
-
pwskill["pathway-enrichment skill"] -.-> enrich
|
|
51
|
-
vizskill["scientific-visualization skill"] -.-> fig
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
## Two Upstream Paths — Pick One
|
|
55
|
-
|
|
56
|
-
The reads → counts stage can be run two ways. They produce equivalent gene counts; choose by context, then stay on that path.
|
|
57
|
-
|
|
58
|
-
| Use **Path A — `nf-core/rnaseq`** when… | Use **Path B — standalone tools** when… |
|
|
59
|
-
|------------------------------------------|------------------------------------------|
|
|
60
|
-
| You want the field-standard, audited, citable pipeline with one command | You have a few samples and want to learn/inspect each step |
|
|
61
|
-
| Many samples, or you'll scale to HPC/cloud | No Nextflow/containers available, or a constrained environment |
|
|
62
|
-
| Reproducibility and a full MultiQC report matter most | You need a non-standard step the pipeline doesn't expose |
|
|
63
|
-
| → Drive it through the **`nextflow`** skill | → Follow `references/upstream-manual.md` |
|
|
64
|
-
|
|
65
|
-
When unsure, prefer **Path A**: `nf-core/rnaseq` already wires together FastQC → trimming → STAR/Salmon → quantification → tximport → MultiQC with sensible, reviewed defaults, which is the most defensible option. Path B exists for transparency and constrained setups.
|
|
66
|
-
|
|
67
|
-
Both paths converge on a **gene-level counts matrix**, after which the workflow is identical.
|
|
68
|
-
|
|
69
|
-
## Setup
|
|
70
|
-
|
|
71
|
-
```bash
|
|
72
|
-
# This skill's glue (bridge + handoffs) — Python
|
|
73
|
-
uv pip install pytximport pandas
|
|
74
|
-
|
|
75
|
-
# Downstream skills install their own deps:
|
|
76
|
-
# pydeseq2 skill -> uv pip install pydeseq2
|
|
77
|
-
# pathway-enrichment skill -> uv pip install gseapy gprofiler-official
|
|
78
|
-
|
|
79
|
-
# Path A (nf-core): only Nextflow + a container engine are needed — see the `nextflow` skill.
|
|
80
|
-
|
|
81
|
-
# Path B (standalone tools): install via bioconda. Pin versions for reproducibility.
|
|
82
|
-
conda create -n rnaseq -c bioconda -c conda-forge \
|
|
83
|
-
fastqc fastp trim-galore "star=2.7.11b" "salmon=1.10.3" subread multiqc
|
|
84
|
-
```
|
|
85
|
-
|
|
86
|
-
Record the exact versions you use (pipeline revision, tool versions, reference genome + annotation release) — they belong in the methods section and make the analysis reproducible.
|
|
87
|
-
|
|
88
|
-
## Quick Start
|
|
89
|
-
|
|
90
|
-
### Path A — nf-core/rnaseq (recommended)
|
|
91
|
-
|
|
92
|
-
```bash
|
|
93
|
-
# 0. Validate the samplesheet first (catches the most common failures early)
|
|
94
|
-
python scripts/validate_samplesheet.py --samplesheet samplesheet.csv
|
|
95
|
-
|
|
96
|
-
# 1. Smoke-test the environment with tiny bundled data
|
|
97
|
-
nextflow run nf-core/rnaseq -r 3.26.0 -profile test,docker --outdir test_results
|
|
98
|
-
|
|
99
|
-
# 2. Real run: pin the revision, pick an aligner, pass a samplesheet + reference
|
|
100
|
-
nextflow run nf-core/rnaseq -r 3.26.0 \
|
|
101
|
-
-profile docker \
|
|
102
|
-
--input samplesheet.csv \
|
|
103
|
-
--genome GRCh38 \
|
|
104
|
-
--aligner star_salmon \
|
|
105
|
-
--outdir results \
|
|
106
|
-
-resume
|
|
107
|
-
```
|
|
108
|
-
|
|
109
|
-
`nf-core/rnaseq` runs tximport internally, so gene counts come out **already merged** — no bridge script needed. Use `results/star_salmon/salmon.merged.gene_counts_length_scaled.tsv` for DE. Samplesheet format, aligner choice, and outputs: `references/upstream-nfcore.md`. For engine/HPC/cloud/container detail, use the **`nextflow`** skill.
|
|
110
|
-
|
|
111
|
-
### Path B — standalone STAR/Salmon (abbreviated)
|
|
112
|
-
|
|
113
|
-
```bash
|
|
114
|
-
fastqc -o qc/ reads/*.fastq.gz # 1. QC raw reads
|
|
115
|
-
fastp -i s1_R1.fq.gz -I s1_R2.fq.gz \
|
|
116
|
-
-o s1_R1.trim.fq.gz -O s1_R2.trim.fq.gz \
|
|
117
|
-
--thread 4 -j s1.fastp.json # 2. Trim adapters/low-quality
|
|
118
|
-
salmon quant -i salmon_index -l A \
|
|
119
|
-
-1 s1_R1.trim.fq.gz -2 s1_R2.trim.fq.gz \
|
|
120
|
-
--gcBias --seqBias -p 8 -o quant/s1 # 3. Quantify (per sample)
|
|
121
|
-
```
|
|
122
|
-
|
|
123
|
-
Full recipes (FastQC, fastp/Trim Galore, STAR index+align+`--quantMode GeneCounts`, Salmon decoy-aware index, featureCounts, strandedness): `references/upstream-manual.md`.
|
|
124
|
-
|
|
125
|
-
### Counts → DE → enrichment (both paths)
|
|
126
|
-
|
|
127
|
-
```bash
|
|
128
|
-
# Path B only: assemble a gene x sample counts matrix + metadata template for PyDESeq2
|
|
129
|
-
python scripts/build_counts_matrix.py --from salmon \
|
|
130
|
-
--quant-dir quant/ --tx2gene tx2gene.tsv --output-dir counts/
|
|
131
|
-
|
|
132
|
-
# Then hand off (see the dedicated skills):
|
|
133
|
-
# pydeseq2: counts.csv + metadata.csv -> DE table (log2FC, padj, stat)
|
|
134
|
-
# pathway-enrichment: rank by `stat` (GSEA) or padj+|LFC| hit list (ORA)
|
|
135
|
-
# scientific-visualization / matplotlib: volcano, MA, heatmap, PCA, enrichment dotplot
|
|
136
|
-
```
|
|
137
|
-
|
|
138
|
-
## Stage-by-Stage Workflow
|
|
139
|
-
|
|
140
|
-
Work top to bottom. Each stage names the skill or file that owns the detail. Don't skip the design/QC stages — they are where bulk RNA-seq studies most often go wrong.
|
|
141
|
-
|
|
142
|
-
1. **Design & sample sheet.** Confirm ≥3 biological replicates per group, identify batch/confounders, and choose the comparison(s). Build the samplesheet and validate it with `scripts/validate_samplesheet.py`. Rationale and rules: `references/design-and-qc.md`.
|
|
143
|
-
2. **Raw-read QC.** FastQC per file; aggregate with MultiQC. Check per-base quality, adapter content, duplication, and over-representation. Thresholds: `references/design-and-qc.md`.
|
|
144
|
-
3. **Trimming.** Remove adapters and low-quality tails (via `fastp` or `Trim Galore`). Re-run FastQC to confirm. Recipes: `references/upstream-manual.md` (Path A does this for you).
|
|
145
|
-
4. **Align / quantify.** STAR (genome alignment + `--quantMode GeneCounts`) and/or Salmon (transcript quasi-mapping, decoy-aware). Determine strandedness — it is easy to get wrong and silently halves your counts. Detail: `references/upstream-manual.md`; pipeline params: `references/upstream-nfcore.md`.
|
|
146
|
-
5. **Build the counts matrix.** Turn quant output into a gene × sample integer matrix and a metadata template (`scripts/build_counts_matrix.py`). The estimated-count and gene-ID-mapping nuances live in `references/counts-and-handoff.md`.
|
|
147
|
-
6. **Differential expression → `pydeseq2` skill.** Load `counts.csv` + `metadata.csv`, set the design (e.g. `~batch + condition`), fit, and test with FDR control. Inspect the PCA and p-value histogram as QC.
|
|
148
|
-
7. **Enrichment → `pathway-enrichment` skill.** For GSEA, rank the *full* gene list by the DESeq2 `stat`; for ORA, pass the thresholded hit list (padj < 0.05, optionally |log2FC| > 1). Map gene IDs to symbols first.
|
|
149
|
-
8. **Figures → `scientific-visualization` skill.** Volcano, MA, sample-distance heatmap, PCA, and enrichment dotplots, plus the MultiQC report for the QC narrative.
|
|
150
|
-
|
|
151
|
-
## The counts → DE bridge (the key glue)
|
|
152
|
-
|
|
153
|
-
This is the one stage with no upstream/downstream skill, so this skill owns it. `scripts/build_counts_matrix.py` converts quant output into exactly what `pydeseq2` expects:
|
|
154
|
-
|
|
155
|
-
- **Salmon** (`--from salmon`): aggregates per-sample `quant.sf` to gene level with `pytximport` using `counts_from_abundance="length_scaled_tpm"` (the right choice for gene-level DE), needs a `tx2gene` map.
|
|
156
|
-
- **STAR** (`--from star`): reads each `ReadsPerGene.out.tab`, selecting the column for your `--strandedness` (unstranded/forward/reverse).
|
|
157
|
-
- **featureCounts** (`--from featurecounts`): parses the combined `featureCounts` matrix.
|
|
158
|
-
|
|
159
|
-
It writes `counts.csv` (genes × samples, integers) and `metadata_template.csv` (one row per sample) for you to fill in. **Salmon/RSEM counts are estimates (non-integer); they are rounded to integers** because PyDESeq2 requires integer counts — see `references/counts-and-handoff.md` for why this is acceptable with `length_scaled_tpm` and how it differs from the offset-based DESeq2+tximport route. That reference also covers Ensembl→symbol mapping (needed before enrichment) and the exact orientation PyDESeq2 wants.
|
|
160
|
-
|
|
161
|
-
## Common Pitfalls
|
|
162
|
-
|
|
163
|
-
These cause most wrong or irreproducible bulk RNA-seq results:
|
|
164
|
-
|
|
165
|
-
1. **Too few replicates.** <3 biological replicates per group gives almost no power and unstable dispersion estimates. More replicates beat deeper sequencing.
|
|
166
|
-
2. **Confounded batch and condition.** If every treated sample was processed on a different day/lane than controls, the effect is unrecoverable. Randomize, and model known batches (`~batch + condition`). See `references/design-and-qc.md`.
|
|
167
|
-
3. **Wrong strandedness.** Choosing the wrong STAR column or featureCounts `-s`/Salmon library type silently discards ~half the reads. Use Salmon `-l A` or infer strandedness, and verify the assigned-reads fraction.
|
|
168
|
-
4. **Feeding TPM/FPKM to DESeq2.** DESeq2 needs raw (or length-scaled) **counts**, never TPM/FPKM/normalized values. The bridge handles this.
|
|
169
|
-
5. **Non-integer counts.** PyDESeq2 requires integers; round Salmon estimates (the bridge does this).
|
|
170
|
-
6. **Gene-ID mismatch into enrichment.** DESeq2 output is often Ensembl IDs; Enrichr/MSigDB want symbols. Map IDs before `pathway-enrichment` or "nothing is significant".
|
|
171
|
-
7. **Skipping post-quant QC.** Always look at the PCA and sample-distance heatmap before trusting DE — they expose swapped labels, outliers, and hidden batches.
|
|
172
|
-
8. **Mixing aligners across samples.** Quantify every sample with the same tool, version, reference, and parameters.
|
|
173
|
-
9. **Unpinned versions.** "latest" pipelines/genomes make results unreproducible; pin `-r`, tool versions, and the genome/annotation release.
|
|
174
|
-
|
|
175
|
-
## Integration with Other Skills
|
|
176
|
-
|
|
177
|
-
- **Upstream execution:** `nextflow` (runs `nf-core/rnaseq`, Path A; HPC/cloud/containers).
|
|
178
|
-
- **Reference data / gene IDs:** `gget` (`gget ref` for genome+GTF, `gget info`/`gget search` for ID mapping), `database-lookup` (Ensembl/NCBI), `biopython`/`pysam` (FASTA/BAM handling).
|
|
179
|
-
- **Differential expression:** `pydeseq2` (the DE engine this skill hands counts to).
|
|
180
|
-
- **Enrichment:** `pathway-enrichment` (ORA + GSEA; its `scripts/run_enrichment.py` reads a DESeq2 results CSV directly).
|
|
181
|
-
- **Figures & reporting:** `scientific-visualization`, `matplotlib`, `seaborn`; `scientific-writing` for the methods/results narrative.
|
|
182
|
-
- **Related but distinct:** `scanpy` (single-cell), `statistical-analysis` (multiple-testing depth).
|
|
183
|
-
|
|
184
|
-
## Reference Files
|
|
185
|
-
|
|
186
|
-
Read the relevant file when you need depth — each is self-contained:
|
|
187
|
-
|
|
188
|
-
- `references/upstream-nfcore.md` — Path A: samplesheet format, `--aligner`/`--pseudo_aligner` choice, key params, the `salmon.merged.gene_counts*.tsv` outputs, MultiQC, and what to hand to `pydeseq2`.
|
|
189
|
-
- `references/upstream-manual.md` — Path B: FastQC, fastp/Trim Galore, STAR genome index + alignment + `--quantMode GeneCounts`, Salmon decoy-aware index + `quant`, featureCounts, and how to determine strandedness.
|
|
190
|
-
- `references/counts-and-handoff.md` — turning quant output into PyDESeq2-ready `counts.csv`/`metadata.csv` (pytximport, STAR column selection, featureCounts), the integer/estimated-count nuance, Ensembl→symbol mapping, and the DE→enrichment rank/hit-list recipe.
|
|
191
|
-
- `references/design-and-qc.md` — experimental design (replication, batch, confounding, design formulas) and QC-metric interpretation (mapping rate, duplication, rRNA, complexity, PCA/outliers) — the defensible-pipeline backbone.
|
|
192
|
-
|
|
193
|
-
## Resources
|
|
194
|
-
|
|
195
|
-
- nf-core/rnaseq: https://nf-co.re/rnaseq · STAR: https://github.com/alexdobin/STAR · Salmon: https://salmon.readthedocs.io
|
|
196
|
-
- fastp: https://github.com/OpenGene/fastp · Trim Galore: https://github.com/FelixKrueger/TrimGalore · MultiQC: https://multiqc.info
|
|
197
|
-
- pytximport: https://pytximport.complextissue.com · featureCounts (Subread): https://subread.sourceforge.net
|
|
198
|
-
- Method background: Love et al. 2014 (DESeq2) DOI 10.1186/s13059-014-0550-8 · Soneson et al. 2015 (tximport) DOI 10.12688/f1000research.7563.2
|