@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,194 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: pathway-enrichment
|
|
3
|
-
description: Run pathway and gene-set enrichment analysis on gene lists or ranked gene data, then interpret the results. Use whenever the user has a set of genes (differentially expressed genes from PyDESeq2/Scanpy, CRISPR-screen hits, cluster marker genes, proteomics hits) and wants to know which biological pathways, GO terms, or gene sets are over-represented or enriched. Covers over-representation analysis (ORA / Enrichr / Fisher / hypergeometric), ranked Gene Set Enrichment Analysis (GSEA / preranked), single-sample scoring (ssGSEA/GSVA), and functional profiling via gseapy, g:Profiler, Enrichr libraries, MSigDB, GO, KEGG, Reactome, and WikiPathways — plus gene-ID mapping, choosing the right background universe, multiple-testing correction, redundancy reduction, dotplots/enrichment maps, and publication-ready tables. Use this for "pathway analysis", "enrichment analysis", "GO enrichment", "KEGG/Reactome pathways", "GSEA", "over-representation", "functional annotation", or "what pathways are my genes in".
|
|
4
|
-
license: MIT
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.0"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Pathway Enrichment
|
|
11
|
-
|
|
12
|
-
## Overview
|
|
13
|
-
|
|
14
|
-
Enrichment analysis answers "what biology is over-represented in my genes?" It is the standard last step after differential expression, a screen, or clustering. There are two core methods, and choosing correctly is the single most important decision:
|
|
15
|
-
|
|
16
|
-
- **ORA (over-representation analysis)** — take a *thresholded* gene list (e.g., padj < 0.05) and test which gene sets it overlaps more than chance, using Fisher's exact / hypergeometric tests. Tools: Enrichr, g:Profiler.
|
|
17
|
-
- **GSEA (gene set enrichment analysis)** — take the *whole ranked list* of genes (no threshold) and test whether each gene set is concentrated toward the top or bottom. Preranked GSEA uses a per-gene score (e.g., the DESeq2 `stat`). Better when effects are broad and subtle.
|
|
18
|
-
|
|
19
|
-
This skill orchestrates these analyses, the gene-set databases behind them, and the interpretation pitfalls that make results wrong or unpublishable.
|
|
20
|
-
|
|
21
|
-
## When to Use This Skill
|
|
22
|
-
|
|
23
|
-
Use this skill when the user wants to:
|
|
24
|
-
- Find enriched GO terms / KEGG / Reactome / WikiPathways / MSigDB Hallmark sets in a gene list.
|
|
25
|
-
- Run GSEA / preranked GSEA on DESeq2, edgeR, limma, or Scanpy `rank_genes_groups` output.
|
|
26
|
-
- Score pathway activity per sample/cell (ssGSEA, GSVA).
|
|
27
|
-
- Interpret, deduplicate, and visualize enrichment results, or build a publication table/figure.
|
|
28
|
-
- Decide between ORA and GSEA, pick gene-set libraries, choose a background, or fix gene-ID problems.
|
|
29
|
-
|
|
30
|
-
For quick one-off Enrichr lookups the `gget` skill (`gget enrichr`) is lighter weight; for raw pathway/interaction APIs (Reactome, KEGG, STRING) see the `database-lookup` skill. Use **this** skill for full, defensible enrichment workflows.
|
|
31
|
-
|
|
32
|
-
## Choosing the Right Method
|
|
33
|
-
|
|
34
|
-
| Situation | Method | Tool / entry point |
|
|
35
|
-
|-----------|--------|--------------------|
|
|
36
|
-
| You have a discrete hit list (DE genes, screen hits, cluster markers) | **ORA** | `gp.enrichr(...)` or g:Profiler |
|
|
37
|
-
| You have a full ranked list (every tested gene + a score) | **Preranked GSEA** | `gp.prerank(...)` |
|
|
38
|
-
| You have an expression matrix + class labels | **GSEA** | `gp.gsea(...)` |
|
|
39
|
-
| You want a pathway score per sample/cell | **ssGSEA / GSVA** | `gp.ssgsea(...)`, `gp.gsva(...)` |
|
|
40
|
-
| You need a custom background or 500+ organisms | **ORA with custom domain** | g:Profiler (`domain_scope='custom'`) |
|
|
41
|
-
| You want TF / signaling *activity* (PROGENy, DoRothEA) | activity inference | see `references/databases-and-gene-sets.md` (decoupler) |
|
|
42
|
-
|
|
43
|
-
When in doubt: a thresholded list → ORA; a ranked table with scores → GSEA. Never threshold a list and then feed it to GSEA — that discards the ranking GSEA depends on.
|
|
44
|
-
|
|
45
|
-
## Setup
|
|
46
|
-
|
|
47
|
-
```bash
|
|
48
|
-
uv pip install gseapy gprofiler-official
|
|
49
|
-
# gseapy pulls pandas, numpy, scipy, matplotlib. Network access is needed for
|
|
50
|
-
# Enrichr, g:Profiler, and MSigDB downloads. For fully offline ORA, use a local
|
|
51
|
-
# GMT file with gp.enrich() (see references/gseapy.md).
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
Verify and list available gene-set libraries (names change over time — never hardcode blindly):
|
|
55
|
-
|
|
56
|
-
```python
|
|
57
|
-
import gseapy as gp
|
|
58
|
-
names = gp.get_library_name(organism="human") # 200+ Enrichr libraries
|
|
59
|
-
print([n for n in names if "Reactome" in n or "KEGG" in n or "Hallmark" in n])
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
## Quick Start
|
|
63
|
-
|
|
64
|
-
### ORA on a hit list (gseapy + Enrichr)
|
|
65
|
-
|
|
66
|
-
```python
|
|
67
|
-
import gseapy as gp
|
|
68
|
-
|
|
69
|
-
# Enrichr libraries expect HGNC gene SYMBOLS (human: UPPERCASE). Map IDs first if needed.
|
|
70
|
-
genes = [g.strip() for g in open("deg_symbols.txt") if g.strip()]
|
|
71
|
-
|
|
72
|
-
enr = gp.enrichr(
|
|
73
|
-
gene_list=genes,
|
|
74
|
-
gene_sets=["MSigDB_Hallmark_2020", "GO_Biological_Process_2023",
|
|
75
|
-
"KEGG_2021_Human", "Reactome_2022"],
|
|
76
|
-
organism="human",
|
|
77
|
-
outdir=None, # in-memory; set a path to also write tables/plots
|
|
78
|
-
)
|
|
79
|
-
res = enr.results
|
|
80
|
-
sig = res[res["Adjusted P-value"] < 0.05].sort_values("Adjusted P-value")
|
|
81
|
-
print(sig[["Gene_set", "Term", "Overlap", "Adjusted P-value", "Combined Score", "Genes"]].head(20))
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
### Preranked GSEA from DESeq2 results
|
|
85
|
-
|
|
86
|
-
```python
|
|
87
|
-
import gseapy as gp
|
|
88
|
-
import pandas as pd
|
|
89
|
-
|
|
90
|
-
res = pd.read_csv("deseq2_results.csv", index_col=0) # index = gene symbols
|
|
91
|
-
# Rank by the test statistic (sign = direction, magnitude = evidence). This is
|
|
92
|
-
# more stable than ranking by log2FoldChange, which is noisy for low-count genes.
|
|
93
|
-
rnk = res["stat"].dropna().sort_values(ascending=False)
|
|
94
|
-
rnk.index = rnk.index.str.upper()
|
|
95
|
-
rnk = rnk[~rnk.index.duplicated(keep="first")]
|
|
96
|
-
|
|
97
|
-
pre = gp.prerank(
|
|
98
|
-
rnk=rnk,
|
|
99
|
-
gene_sets=["MSigDB_Hallmark_2020", "GO_Biological_Process_2023"],
|
|
100
|
-
min_size=15, max_size=500, # drop tiny/huge sets (noisy or generic)
|
|
101
|
-
permutation_num=1000, seed=123, # seed = reproducible p-values
|
|
102
|
-
threads=4, outdir=None,
|
|
103
|
-
)
|
|
104
|
-
out = pre.res2d.sort_values("FDR q-val")
|
|
105
|
-
print(out[["Term", "ES", "NES", "NOM p-val", "FDR q-val", "Lead_genes"]].head(20))
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
If you have no `stat` column, build the rank from `sign(log2FoldChange) * -log10(pvalue)`.
|
|
109
|
-
|
|
110
|
-
## Core Workflow
|
|
111
|
-
|
|
112
|
-
For a defensible analysis, work through these steps. The middle steps (ID type, background) are where results most often silently go wrong.
|
|
113
|
-
|
|
114
|
-
### Step 1 — Pin down inputs and pick the method
|
|
115
|
-
Confirm: which genes, what organism, is there a per-gene score (→ GSEA) or just a list (→ ORA), and what comparison they represent (direction matters for interpretation).
|
|
116
|
-
|
|
117
|
-
### Step 2 — Get gene IDs into the right namespace
|
|
118
|
-
Enrichr/MSigDB libraries are keyed by **gene symbols** (human UPPERCASE, mouse Title-case). If you have Ensembl/Entrez IDs, convert first. See `references/databases-and-gene-sets.md` for `gp.Biomart`, g:Profiler `g:Convert`, and `mygene`. A silent ID mismatch is the #1 cause of "nothing is significant".
|
|
119
|
-
|
|
120
|
-
### Step 3 — Choose gene-set libraries to match the question
|
|
121
|
-
Hallmark (broad themes) → GO:BP (mechanism) → KEGG/Reactome/WikiPathways (curated pathways) → C7 (immune), etc. Don't run 50 libraries; pick 2–4 that fit the biology. Catalog and selection guidance: `references/databases-and-gene-sets.md`.
|
|
122
|
-
|
|
123
|
-
### Step 4 — Set the background universe (ORA only)
|
|
124
|
-
The background must be the genes that *could* have been detected in your assay (e.g., all expressed/tested genes), not the whole genome. The wrong background inflates significance. Enrichr uses a fixed background; when background matters, use g:Profiler with `domain_scope='custom'` + your `background`, or `gp.enrich()` with an explicit background. Rationale in `references/interpretation.md`.
|
|
125
|
-
|
|
126
|
-
### Step 5 — Run the analysis
|
|
127
|
-
Use the Quick Start patterns or the bundled `scripts/run_enrichment.py`. For GSEA always set a `seed` and report `permutation_num`.
|
|
128
|
-
|
|
129
|
-
### Step 6 — Filter on adjusted p-values
|
|
130
|
-
Use `Adjusted P-value` (ORA, Benjamini–Hochberg) or `FDR q-val` (GSEA), not raw p-values. Typical cutoff 0.05; also check the overlap/gene count so a "hit" isn't 1 gene out of a 2000-gene set.
|
|
131
|
-
|
|
132
|
-
### Step 7 — Visualize
|
|
133
|
-
Dotplots, bar plots, enrichment maps, and GSEA running-score plots are built into gseapy (`gp.dotplot`, `gp.barplot`, `gp.enrichment_map`, `gp.gseaplot`). See `references/gseapy.md`.
|
|
134
|
-
|
|
135
|
-
### Step 8 — Reduce redundancy and interpret
|
|
136
|
-
GO especially returns many near-duplicate terms. Collapse with an enrichment map (term–term similarity), leading-edge overlap, or parent terms, and report representative terms. Interpretation framework and a publication-table format are in `references/interpretation.md`.
|
|
137
|
-
|
|
138
|
-
## Helper Script
|
|
139
|
-
|
|
140
|
-
`scripts/run_enrichment.py` runs ORA or GSEA end-to-end and writes a results table plus a dotplot, handling the boilerplate (symbol cleanup, dedup, NA removal, rank construction from a DESeq2 table, per-library FDR filtering).
|
|
141
|
-
|
|
142
|
-
```bash
|
|
143
|
-
# ORA from a hit list (one gene symbol per line)
|
|
144
|
-
python scripts/run_enrichment.py ora \
|
|
145
|
-
--genes deg_symbols.txt \
|
|
146
|
-
--libraries MSigDB_Hallmark_2020 GO_Biological_Process_2023 KEGG_2021_Human \
|
|
147
|
-
--organism human --outdir results/
|
|
148
|
-
|
|
149
|
-
# Preranked GSEA from a DESeq2 results CSV (auto-builds the rank from `stat`)
|
|
150
|
-
python scripts/run_enrichment.py gsea \
|
|
151
|
-
--deseq2 deseq2_results.csv \
|
|
152
|
-
--libraries MSigDB_Hallmark_2020 GO_Biological_Process_2023 \
|
|
153
|
-
--organism human --outdir results/ --seed 123
|
|
154
|
-
|
|
155
|
-
# Preranked GSEA from an explicit 2-column rank file (gene,score)
|
|
156
|
-
python scripts/run_enrichment.py gsea --rnk ranked_genes.csv --outdir results/
|
|
157
|
-
```
|
|
158
|
-
|
|
159
|
-
Run `python scripts/run_enrichment.py --help` for all options (background file, FDR cutoff, min/max set size, permutations).
|
|
160
|
-
|
|
161
|
-
## Common Pitfalls
|
|
162
|
-
|
|
163
|
-
These cause most wrong or irreproducible results:
|
|
164
|
-
|
|
165
|
-
1. **Gene-ID / organism mismatch** — symbols vs Ensembl, human vs mouse casing. Map IDs and set `organism` correctly, or matches silently drop to ~zero.
|
|
166
|
-
2. **Wrong background (ORA)** — using the whole genome instead of the tested/expressed gene set inflates p-values. Set a custom background when it matters.
|
|
167
|
-
3. **Thresholding before GSEA** — GSEA needs the *full* ranked list; only ORA uses a cut list.
|
|
168
|
-
4. **Ranking GSEA by log2FoldChange alone** — unstable for low-count genes; prefer `stat` or `sign(LFC) * -log10(p)`.
|
|
169
|
-
5. **Multiple-testing across libraries** — FDR is computed *within* a library; running many libraries multiplies tests. Report per-library FDR and stay conservative.
|
|
170
|
-
6. **Redundant GO terms** — don't report 40 variants of the same term; collapse and show representatives.
|
|
171
|
-
7. **Significance ≠ relevance** — check the overlap count and gene-set size; tiny sets reach significance trivially.
|
|
172
|
-
8. **List too short/long for ORA** — <10 genes is underpowered; >2000 loses specificity (consider GSEA instead).
|
|
173
|
-
9. **No reproducibility metadata** — Enrichr/GO libraries are versioned and drift over time. Record library names+date and set a GSEA `seed`.
|
|
174
|
-
|
|
175
|
-
## Integration with Other Skills
|
|
176
|
-
|
|
177
|
-
- **Upstream (where genes come from):** `pydeseq2` (DE genes + `stat` for GSEA), `scanpy` (`rank_genes_groups` markers / scores), `depmap`/`pytdc` (screen hits), proteomics skills (`pyopenms`, `matchms`).
|
|
178
|
-
- **Databases / IDs:** `database-lookup` (Reactome, KEGG, STRING, Gene Ontology APIs), `gget` (`gget enrichr` quick path, `gget info` for ID mapping), `bioservices`.
|
|
179
|
-
- **Downstream:** `scientific-visualization` (custom figures), `networkx` (enrichment-map graphs), `scientific-writing` / `literature-review` (interpret and cite), `statistical-analysis` (multiple-testing details).
|
|
180
|
-
|
|
181
|
-
## Reference Files
|
|
182
|
-
|
|
183
|
-
Read the relevant file when you need depth:
|
|
184
|
-
|
|
185
|
-
- `references/gseapy.md` — full gseapy API: `enrichr`, offline `enrich`, `prerank`, `gsea`, `ssgsea`, `gsva`, `Msigdb`, `Biomart`, `get_library_name`/`read_gmt`, every plot, result-column meanings, GMT/offline usage, and troubleshooting (rate limits, empty results).
|
|
186
|
-
- `references/databases-and-gene-sets.md` — GO, KEGG, Reactome, WikiPathways, MSigDB collections, Enrichr library naming, g:Profiler sources, organism handling, gene-ID conversion, library selection by question, and pointers to Reactome/STRING APIs and decoupler activity inference.
|
|
187
|
-
- `references/interpretation.md` — ORA vs GSEA statistics, background-universe choice, multiple-testing methods (BH vs g:SCS vs Bonferroni), leading-edge genes, redundancy reduction, effect vs significance, a publication-table template, and reproducibility checklist.
|
|
188
|
-
|
|
189
|
-
## Resources
|
|
190
|
-
|
|
191
|
-
- gseapy docs: https://gseapy.readthedocs.io/ · repo: https://github.com/zqfang/GSEApy
|
|
192
|
-
- g:Profiler: https://biit.cs.ut.ee/gprofiler/ · Python client: https://pypi.org/project/gprofiler-official/
|
|
193
|
-
- Enrichr: https://maayanlab.cloud/Enrichr/ · MSigDB: https://www.gsea-msigdb.org/gsea/msigdb/
|
|
194
|
-
- GSEA method: Subramanian et al. (2005) PNAS, DOI: 10.1073/pnas.0506580102
|
package/skills/pdf/SKILL.md
DELETED
|
@@ -1,322 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: pdf
|
|
3
|
-
description: Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.
|
|
4
|
-
license: Proprietary. LICENSE.txt has complete terms
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.2"
|
|
7
|
-
skill-author: Anthropic, PBC
|
|
8
|
-
source: https://github.com/anthropics/skills/tree/main/skills/pdf
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# PDF Processing Guide
|
|
12
|
-
|
|
13
|
-
## Overview
|
|
14
|
-
|
|
15
|
-
This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read forms.md and follow its instructions.
|
|
16
|
-
|
|
17
|
-
## Quick Start
|
|
18
|
-
|
|
19
|
-
```python
|
|
20
|
-
from pypdf import PdfReader, PdfWriter
|
|
21
|
-
|
|
22
|
-
# Read a PDF
|
|
23
|
-
reader = PdfReader("document.pdf")
|
|
24
|
-
print(f"Pages: {len(reader.pages)}")
|
|
25
|
-
|
|
26
|
-
# Extract text
|
|
27
|
-
text = ""
|
|
28
|
-
for page in reader.pages:
|
|
29
|
-
text += page.extract_text()
|
|
30
|
-
```
|
|
31
|
-
|
|
32
|
-
## Python Libraries
|
|
33
|
-
|
|
34
|
-
### pypdf - Basic Operations
|
|
35
|
-
|
|
36
|
-
#### Merge PDFs
|
|
37
|
-
```python
|
|
38
|
-
from pypdf import PdfWriter, PdfReader
|
|
39
|
-
|
|
40
|
-
writer = PdfWriter()
|
|
41
|
-
for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:
|
|
42
|
-
reader = PdfReader(pdf_file)
|
|
43
|
-
for page in reader.pages:
|
|
44
|
-
writer.add_page(page)
|
|
45
|
-
|
|
46
|
-
with open("merged.pdf", "wb") as output:
|
|
47
|
-
writer.write(output)
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
#### Split PDF
|
|
51
|
-
```python
|
|
52
|
-
reader = PdfReader("input.pdf")
|
|
53
|
-
for i, page in enumerate(reader.pages):
|
|
54
|
-
writer = PdfWriter()
|
|
55
|
-
writer.add_page(page)
|
|
56
|
-
with open(f"page_{i+1}.pdf", "wb") as output:
|
|
57
|
-
writer.write(output)
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
#### Extract Metadata
|
|
61
|
-
```python
|
|
62
|
-
reader = PdfReader("document.pdf")
|
|
63
|
-
meta = reader.metadata
|
|
64
|
-
print(f"Title: {meta.title}")
|
|
65
|
-
print(f"Author: {meta.author}")
|
|
66
|
-
print(f"Subject: {meta.subject}")
|
|
67
|
-
print(f"Creator: {meta.creator}")
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
#### Rotate Pages
|
|
71
|
-
```python
|
|
72
|
-
reader = PdfReader("input.pdf")
|
|
73
|
-
writer = PdfWriter()
|
|
74
|
-
|
|
75
|
-
page = reader.pages[0]
|
|
76
|
-
page.rotate(90) # Rotate 90 degrees clockwise
|
|
77
|
-
writer.add_page(page)
|
|
78
|
-
|
|
79
|
-
with open("rotated.pdf", "wb") as output:
|
|
80
|
-
writer.write(output)
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
### pdfplumber - Text and Table Extraction
|
|
84
|
-
|
|
85
|
-
#### Extract Text with Layout
|
|
86
|
-
```python
|
|
87
|
-
import pdfplumber
|
|
88
|
-
|
|
89
|
-
with pdfplumber.open("document.pdf") as pdf:
|
|
90
|
-
for page in pdf.pages:
|
|
91
|
-
text = page.extract_text()
|
|
92
|
-
print(text)
|
|
93
|
-
```
|
|
94
|
-
|
|
95
|
-
#### Extract Tables
|
|
96
|
-
```python
|
|
97
|
-
with pdfplumber.open("document.pdf") as pdf:
|
|
98
|
-
for i, page in enumerate(pdf.pages):
|
|
99
|
-
tables = page.extract_tables()
|
|
100
|
-
for j, table in enumerate(tables):
|
|
101
|
-
print(f"Table {j+1} on page {i+1}:")
|
|
102
|
-
for row in table:
|
|
103
|
-
print(row)
|
|
104
|
-
```
|
|
105
|
-
|
|
106
|
-
#### Advanced Table Extraction
|
|
107
|
-
```python
|
|
108
|
-
import pandas as pd
|
|
109
|
-
|
|
110
|
-
with pdfplumber.open("document.pdf") as pdf:
|
|
111
|
-
all_tables = []
|
|
112
|
-
for page in pdf.pages:
|
|
113
|
-
tables = page.extract_tables()
|
|
114
|
-
for table in tables:
|
|
115
|
-
if table: # Check if table is not empty
|
|
116
|
-
df = pd.DataFrame(table[1:], columns=table[0])
|
|
117
|
-
all_tables.append(df)
|
|
118
|
-
|
|
119
|
-
# Combine all tables
|
|
120
|
-
if all_tables:
|
|
121
|
-
combined_df = pd.concat(all_tables, ignore_index=True)
|
|
122
|
-
combined_df.to_excel("extracted_tables.xlsx", index=False)
|
|
123
|
-
```
|
|
124
|
-
|
|
125
|
-
### reportlab - Create PDFs
|
|
126
|
-
|
|
127
|
-
#### Basic PDF Creation
|
|
128
|
-
```python
|
|
129
|
-
from reportlab.lib.pagesizes import letter
|
|
130
|
-
from reportlab.pdfgen import canvas
|
|
131
|
-
|
|
132
|
-
c = canvas.Canvas("hello.pdf", pagesize=letter)
|
|
133
|
-
width, height = letter
|
|
134
|
-
|
|
135
|
-
# Add text
|
|
136
|
-
c.drawString(100, height - 100, "Hello World!")
|
|
137
|
-
c.drawString(100, height - 120, "This is a PDF created with reportlab")
|
|
138
|
-
|
|
139
|
-
# Add a line
|
|
140
|
-
c.line(100, height - 140, 400, height - 140)
|
|
141
|
-
|
|
142
|
-
# Save
|
|
143
|
-
c.save()
|
|
144
|
-
```
|
|
145
|
-
|
|
146
|
-
#### Create PDF with Multiple Pages
|
|
147
|
-
```python
|
|
148
|
-
from reportlab.lib.pagesizes import letter
|
|
149
|
-
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak
|
|
150
|
-
from reportlab.lib.styles import getSampleStyleSheet
|
|
151
|
-
|
|
152
|
-
doc = SimpleDocTemplate("report.pdf", pagesize=letter)
|
|
153
|
-
styles = getSampleStyleSheet()
|
|
154
|
-
story = []
|
|
155
|
-
|
|
156
|
-
# Add content
|
|
157
|
-
title = Paragraph("Report Title", styles['Title'])
|
|
158
|
-
story.append(title)
|
|
159
|
-
story.append(Spacer(1, 12))
|
|
160
|
-
|
|
161
|
-
body = Paragraph("This is the body of the report. " * 20, styles['Normal'])
|
|
162
|
-
story.append(body)
|
|
163
|
-
story.append(PageBreak())
|
|
164
|
-
|
|
165
|
-
# Page 2
|
|
166
|
-
story.append(Paragraph("Page 2", styles['Heading1']))
|
|
167
|
-
story.append(Paragraph("Content for page 2", styles['Normal']))
|
|
168
|
-
|
|
169
|
-
# Build PDF
|
|
170
|
-
doc.build(story)
|
|
171
|
-
```
|
|
172
|
-
|
|
173
|
-
#### Subscripts and Superscripts
|
|
174
|
-
|
|
175
|
-
**IMPORTANT**: Never use Unicode subscript/superscript characters (₀₁₂₃₄₅₆₇₈₉, ⁰¹²³⁴⁵⁶⁷⁸⁹) in ReportLab PDFs. The built-in fonts do not include these glyphs, causing them to render as solid black boxes.
|
|
176
|
-
|
|
177
|
-
Instead, use ReportLab's XML markup tags in Paragraph objects:
|
|
178
|
-
```python
|
|
179
|
-
from reportlab.platypus import Paragraph
|
|
180
|
-
from reportlab.lib.styles import getSampleStyleSheet
|
|
181
|
-
|
|
182
|
-
styles = getSampleStyleSheet()
|
|
183
|
-
|
|
184
|
-
# Subscripts: use <sub> tag
|
|
185
|
-
chemical = Paragraph("H<sub>2</sub>O", styles['Normal'])
|
|
186
|
-
|
|
187
|
-
# Superscripts: use <super> tag
|
|
188
|
-
squared = Paragraph("x<super>2</super> + y<super>2</super>", styles['Normal'])
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
For canvas-drawn text (not Paragraph objects), manually adjust font the size and position rather than using Unicode subscripts/superscripts.
|
|
192
|
-
|
|
193
|
-
## Command-Line Tools
|
|
194
|
-
|
|
195
|
-
### pdftotext (poppler-utils)
|
|
196
|
-
```bash
|
|
197
|
-
# Extract text
|
|
198
|
-
pdftotext input.pdf output.txt
|
|
199
|
-
|
|
200
|
-
# Extract text preserving layout
|
|
201
|
-
pdftotext -layout input.pdf output.txt
|
|
202
|
-
|
|
203
|
-
# Extract specific pages
|
|
204
|
-
pdftotext -f 1 -l 5 input.pdf output.txt # Pages 1-5
|
|
205
|
-
```
|
|
206
|
-
|
|
207
|
-
### qpdf
|
|
208
|
-
```bash
|
|
209
|
-
# Merge PDFs
|
|
210
|
-
qpdf --empty --pages file1.pdf file2.pdf -- merged.pdf
|
|
211
|
-
|
|
212
|
-
# Split pages
|
|
213
|
-
qpdf input.pdf --pages . 1-5 -- pages1-5.pdf
|
|
214
|
-
qpdf input.pdf --pages . 6-10 -- pages6-10.pdf
|
|
215
|
-
|
|
216
|
-
# Rotate pages
|
|
217
|
-
qpdf input.pdf output.pdf --rotate=+90:1 # Rotate page 1 by 90 degrees
|
|
218
|
-
|
|
219
|
-
# Remove password
|
|
220
|
-
qpdf --password=mypassword --decrypt encrypted.pdf decrypted.pdf
|
|
221
|
-
```
|
|
222
|
-
|
|
223
|
-
### pdftk (if available)
|
|
224
|
-
```bash
|
|
225
|
-
# Merge
|
|
226
|
-
pdftk file1.pdf file2.pdf cat output merged.pdf
|
|
227
|
-
|
|
228
|
-
# Split
|
|
229
|
-
pdftk input.pdf burst
|
|
230
|
-
|
|
231
|
-
# Rotate
|
|
232
|
-
pdftk input.pdf rotate 1east output rotated.pdf
|
|
233
|
-
```
|
|
234
|
-
|
|
235
|
-
## Common Tasks
|
|
236
|
-
|
|
237
|
-
### Extract Text from Scanned PDFs
|
|
238
|
-
```python
|
|
239
|
-
# Requires: uv pip install pytesseract pdf2image
|
|
240
|
-
import pytesseract
|
|
241
|
-
from pdf2image import convert_from_path
|
|
242
|
-
|
|
243
|
-
# Convert PDF to images
|
|
244
|
-
images = convert_from_path('scanned.pdf')
|
|
245
|
-
|
|
246
|
-
# OCR each page
|
|
247
|
-
text = ""
|
|
248
|
-
for i, image in enumerate(images):
|
|
249
|
-
text += f"Page {i+1}:\n"
|
|
250
|
-
text += pytesseract.image_to_string(image)
|
|
251
|
-
text += "\n\n"
|
|
252
|
-
|
|
253
|
-
print(text)
|
|
254
|
-
```
|
|
255
|
-
|
|
256
|
-
### Add Watermark
|
|
257
|
-
```python
|
|
258
|
-
from pypdf import PdfReader, PdfWriter
|
|
259
|
-
|
|
260
|
-
# Create watermark (or load existing)
|
|
261
|
-
watermark = PdfReader("watermark.pdf").pages[0]
|
|
262
|
-
|
|
263
|
-
# Apply to all pages
|
|
264
|
-
reader = PdfReader("document.pdf")
|
|
265
|
-
writer = PdfWriter()
|
|
266
|
-
|
|
267
|
-
for page in reader.pages:
|
|
268
|
-
page.merge_page(watermark)
|
|
269
|
-
writer.add_page(page)
|
|
270
|
-
|
|
271
|
-
with open("watermarked.pdf", "wb") as output:
|
|
272
|
-
writer.write(output)
|
|
273
|
-
```
|
|
274
|
-
|
|
275
|
-
### Extract Images
|
|
276
|
-
```bash
|
|
277
|
-
# Using pdfimages (poppler-utils)
|
|
278
|
-
pdfimages -j input.pdf output_prefix
|
|
279
|
-
|
|
280
|
-
# This extracts all images as output_prefix-000.jpg, output_prefix-001.jpg, etc.
|
|
281
|
-
```
|
|
282
|
-
|
|
283
|
-
### Password Protection
|
|
284
|
-
```python
|
|
285
|
-
from pypdf import PdfReader, PdfWriter
|
|
286
|
-
|
|
287
|
-
reader = PdfReader("input.pdf")
|
|
288
|
-
writer = PdfWriter()
|
|
289
|
-
|
|
290
|
-
for page in reader.pages:
|
|
291
|
-
writer.add_page(page)
|
|
292
|
-
|
|
293
|
-
# Add password
|
|
294
|
-
writer.encrypt("userpassword", "ownerpassword")
|
|
295
|
-
|
|
296
|
-
with open("encrypted.pdf", "wb") as output:
|
|
297
|
-
writer.write(output)
|
|
298
|
-
```
|
|
299
|
-
|
|
300
|
-
## Quick Reference
|
|
301
|
-
|
|
302
|
-
| Task | Best Tool | Command/Code |
|
|
303
|
-
|------|-----------|--------------|
|
|
304
|
-
| Merge PDFs | pypdf | `writer.add_page(page)` |
|
|
305
|
-
| Split PDFs | pypdf | One page per file |
|
|
306
|
-
| Extract text | pdfplumber | `page.extract_text()` |
|
|
307
|
-
| Extract tables | pdfplumber | `page.extract_tables()` |
|
|
308
|
-
| Create PDFs | reportlab | Canvas or Platypus |
|
|
309
|
-
| Command line merge | qpdf | `qpdf --empty --pages ...` |
|
|
310
|
-
| OCR scanned PDFs | pytesseract | Convert to image first |
|
|
311
|
-
| Fill PDF forms | pdf-lib or pypdf (see forms.md) | See forms.md |
|
|
312
|
-
|
|
313
|
-
## Next Steps
|
|
314
|
-
|
|
315
|
-
- For advanced pypdfium2 usage, see reference.md
|
|
316
|
-
- For JavaScript libraries (pdf-lib), see reference.md
|
|
317
|
-
- If you need to fill out a PDF form, follow the instructions in forms.md
|
|
318
|
-
- For troubleshooting guides, see reference.md
|
|
319
|
-
|
|
320
|
-
---
|
|
321
|
-
|
|
322
|
-
*This skill is created and maintained by [Anthropic](https://github.com/anthropics/skills/tree/main/skills/pdf). Vendored here unmodified except for frontmatter metadata and the case of the `reference.md`/`forms.md` links, which upstream writes uppercase; see LICENSE.txt for terms.*
|