@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,194 +0,0 @@
1
- ---
2
- name: pathway-enrichment
3
- description: Run pathway and gene-set enrichment analysis on gene lists or ranked gene data, then interpret the results. Use whenever the user has a set of genes (differentially expressed genes from PyDESeq2/Scanpy, CRISPR-screen hits, cluster marker genes, proteomics hits) and wants to know which biological pathways, GO terms, or gene sets are over-represented or enriched. Covers over-representation analysis (ORA / Enrichr / Fisher / hypergeometric), ranked Gene Set Enrichment Analysis (GSEA / preranked), single-sample scoring (ssGSEA/GSVA), and functional profiling via gseapy, g:Profiler, Enrichr libraries, MSigDB, GO, KEGG, Reactome, and WikiPathways — plus gene-ID mapping, choosing the right background universe, multiple-testing correction, redundancy reduction, dotplots/enrichment maps, and publication-ready tables. Use this for "pathway analysis", "enrichment analysis", "GO enrichment", "KEGG/Reactome pathways", "GSEA", "over-representation", "functional annotation", or "what pathways are my genes in".
4
- license: MIT
5
- metadata:
6
- version: "1.0"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # Pathway Enrichment
11
-
12
- ## Overview
13
-
14
- Enrichment analysis answers "what biology is over-represented in my genes?" It is the standard last step after differential expression, a screen, or clustering. There are two core methods, and choosing correctly is the single most important decision:
15
-
16
- - **ORA (over-representation analysis)** — take a *thresholded* gene list (e.g., padj < 0.05) and test which gene sets it overlaps more than chance, using Fisher's exact / hypergeometric tests. Tools: Enrichr, g:Profiler.
17
- - **GSEA (gene set enrichment analysis)** — take the *whole ranked list* of genes (no threshold) and test whether each gene set is concentrated toward the top or bottom. Preranked GSEA uses a per-gene score (e.g., the DESeq2 `stat`). Better when effects are broad and subtle.
18
-
19
- This skill orchestrates these analyses, the gene-set databases behind them, and the interpretation pitfalls that make results wrong or unpublishable.
20
-
21
- ## When to Use This Skill
22
-
23
- Use this skill when the user wants to:
24
- - Find enriched GO terms / KEGG / Reactome / WikiPathways / MSigDB Hallmark sets in a gene list.
25
- - Run GSEA / preranked GSEA on DESeq2, edgeR, limma, or Scanpy `rank_genes_groups` output.
26
- - Score pathway activity per sample/cell (ssGSEA, GSVA).
27
- - Interpret, deduplicate, and visualize enrichment results, or build a publication table/figure.
28
- - Decide between ORA and GSEA, pick gene-set libraries, choose a background, or fix gene-ID problems.
29
-
30
- For quick one-off Enrichr lookups the `gget` skill (`gget enrichr`) is lighter weight; for raw pathway/interaction APIs (Reactome, KEGG, STRING) see the `database-lookup` skill. Use **this** skill for full, defensible enrichment workflows.
31
-
32
- ## Choosing the Right Method
33
-
34
- | Situation | Method | Tool / entry point |
35
- |-----------|--------|--------------------|
36
- | You have a discrete hit list (DE genes, screen hits, cluster markers) | **ORA** | `gp.enrichr(...)` or g:Profiler |
37
- | You have a full ranked list (every tested gene + a score) | **Preranked GSEA** | `gp.prerank(...)` |
38
- | You have an expression matrix + class labels | **GSEA** | `gp.gsea(...)` |
39
- | You want a pathway score per sample/cell | **ssGSEA / GSVA** | `gp.ssgsea(...)`, `gp.gsva(...)` |
40
- | You need a custom background or 500+ organisms | **ORA with custom domain** | g:Profiler (`domain_scope='custom'`) |
41
- | You want TF / signaling *activity* (PROGENy, DoRothEA) | activity inference | see `references/databases-and-gene-sets.md` (decoupler) |
42
-
43
- When in doubt: a thresholded list → ORA; a ranked table with scores → GSEA. Never threshold a list and then feed it to GSEA — that discards the ranking GSEA depends on.
44
-
45
- ## Setup
46
-
47
- ```bash
48
- uv pip install gseapy gprofiler-official
49
- # gseapy pulls pandas, numpy, scipy, matplotlib. Network access is needed for
50
- # Enrichr, g:Profiler, and MSigDB downloads. For fully offline ORA, use a local
51
- # GMT file with gp.enrich() (see references/gseapy.md).
52
- ```
53
-
54
- Verify and list available gene-set libraries (names change over time — never hardcode blindly):
55
-
56
- ```python
57
- import gseapy as gp
58
- names = gp.get_library_name(organism="human") # 200+ Enrichr libraries
59
- print([n for n in names if "Reactome" in n or "KEGG" in n or "Hallmark" in n])
60
- ```
61
-
62
- ## Quick Start
63
-
64
- ### ORA on a hit list (gseapy + Enrichr)
65
-
66
- ```python
67
- import gseapy as gp
68
-
69
- # Enrichr libraries expect HGNC gene SYMBOLS (human: UPPERCASE). Map IDs first if needed.
70
- genes = [g.strip() for g in open("deg_symbols.txt") if g.strip()]
71
-
72
- enr = gp.enrichr(
73
- gene_list=genes,
74
- gene_sets=["MSigDB_Hallmark_2020", "GO_Biological_Process_2023",
75
- "KEGG_2021_Human", "Reactome_2022"],
76
- organism="human",
77
- outdir=None, # in-memory; set a path to also write tables/plots
78
- )
79
- res = enr.results
80
- sig = res[res["Adjusted P-value"] < 0.05].sort_values("Adjusted P-value")
81
- print(sig[["Gene_set", "Term", "Overlap", "Adjusted P-value", "Combined Score", "Genes"]].head(20))
82
- ```
83
-
84
- ### Preranked GSEA from DESeq2 results
85
-
86
- ```python
87
- import gseapy as gp
88
- import pandas as pd
89
-
90
- res = pd.read_csv("deseq2_results.csv", index_col=0) # index = gene symbols
91
- # Rank by the test statistic (sign = direction, magnitude = evidence). This is
92
- # more stable than ranking by log2FoldChange, which is noisy for low-count genes.
93
- rnk = res["stat"].dropna().sort_values(ascending=False)
94
- rnk.index = rnk.index.str.upper()
95
- rnk = rnk[~rnk.index.duplicated(keep="first")]
96
-
97
- pre = gp.prerank(
98
- rnk=rnk,
99
- gene_sets=["MSigDB_Hallmark_2020", "GO_Biological_Process_2023"],
100
- min_size=15, max_size=500, # drop tiny/huge sets (noisy or generic)
101
- permutation_num=1000, seed=123, # seed = reproducible p-values
102
- threads=4, outdir=None,
103
- )
104
- out = pre.res2d.sort_values("FDR q-val")
105
- print(out[["Term", "ES", "NES", "NOM p-val", "FDR q-val", "Lead_genes"]].head(20))
106
- ```
107
-
108
- If you have no `stat` column, build the rank from `sign(log2FoldChange) * -log10(pvalue)`.
109
-
110
- ## Core Workflow
111
-
112
- For a defensible analysis, work through these steps. The middle steps (ID type, background) are where results most often silently go wrong.
113
-
114
- ### Step 1 — Pin down inputs and pick the method
115
- Confirm: which genes, what organism, is there a per-gene score (→ GSEA) or just a list (→ ORA), and what comparison they represent (direction matters for interpretation).
116
-
117
- ### Step 2 — Get gene IDs into the right namespace
118
- Enrichr/MSigDB libraries are keyed by **gene symbols** (human UPPERCASE, mouse Title-case). If you have Ensembl/Entrez IDs, convert first. See `references/databases-and-gene-sets.md` for `gp.Biomart`, g:Profiler `g:Convert`, and `mygene`. A silent ID mismatch is the #1 cause of "nothing is significant".
119
-
120
- ### Step 3 — Choose gene-set libraries to match the question
121
- Hallmark (broad themes) → GO:BP (mechanism) → KEGG/Reactome/WikiPathways (curated pathways) → C7 (immune), etc. Don't run 50 libraries; pick 2–4 that fit the biology. Catalog and selection guidance: `references/databases-and-gene-sets.md`.
122
-
123
- ### Step 4 — Set the background universe (ORA only)
124
- The background must be the genes that *could* have been detected in your assay (e.g., all expressed/tested genes), not the whole genome. The wrong background inflates significance. Enrichr uses a fixed background; when background matters, use g:Profiler with `domain_scope='custom'` + your `background`, or `gp.enrich()` with an explicit background. Rationale in `references/interpretation.md`.
125
-
126
- ### Step 5 — Run the analysis
127
- Use the Quick Start patterns or the bundled `scripts/run_enrichment.py`. For GSEA always set a `seed` and report `permutation_num`.
128
-
129
- ### Step 6 — Filter on adjusted p-values
130
- Use `Adjusted P-value` (ORA, Benjamini–Hochberg) or `FDR q-val` (GSEA), not raw p-values. Typical cutoff 0.05; also check the overlap/gene count so a "hit" isn't 1 gene out of a 2000-gene set.
131
-
132
- ### Step 7 — Visualize
133
- Dotplots, bar plots, enrichment maps, and GSEA running-score plots are built into gseapy (`gp.dotplot`, `gp.barplot`, `gp.enrichment_map`, `gp.gseaplot`). See `references/gseapy.md`.
134
-
135
- ### Step 8 — Reduce redundancy and interpret
136
- GO especially returns many near-duplicate terms. Collapse with an enrichment map (term–term similarity), leading-edge overlap, or parent terms, and report representative terms. Interpretation framework and a publication-table format are in `references/interpretation.md`.
137
-
138
- ## Helper Script
139
-
140
- `scripts/run_enrichment.py` runs ORA or GSEA end-to-end and writes a results table plus a dotplot, handling the boilerplate (symbol cleanup, dedup, NA removal, rank construction from a DESeq2 table, per-library FDR filtering).
141
-
142
- ```bash
143
- # ORA from a hit list (one gene symbol per line)
144
- python scripts/run_enrichment.py ora \
145
- --genes deg_symbols.txt \
146
- --libraries MSigDB_Hallmark_2020 GO_Biological_Process_2023 KEGG_2021_Human \
147
- --organism human --outdir results/
148
-
149
- # Preranked GSEA from a DESeq2 results CSV (auto-builds the rank from `stat`)
150
- python scripts/run_enrichment.py gsea \
151
- --deseq2 deseq2_results.csv \
152
- --libraries MSigDB_Hallmark_2020 GO_Biological_Process_2023 \
153
- --organism human --outdir results/ --seed 123
154
-
155
- # Preranked GSEA from an explicit 2-column rank file (gene,score)
156
- python scripts/run_enrichment.py gsea --rnk ranked_genes.csv --outdir results/
157
- ```
158
-
159
- Run `python scripts/run_enrichment.py --help` for all options (background file, FDR cutoff, min/max set size, permutations).
160
-
161
- ## Common Pitfalls
162
-
163
- These cause most wrong or irreproducible results:
164
-
165
- 1. **Gene-ID / organism mismatch** — symbols vs Ensembl, human vs mouse casing. Map IDs and set `organism` correctly, or matches silently drop to ~zero.
166
- 2. **Wrong background (ORA)** — using the whole genome instead of the tested/expressed gene set inflates p-values. Set a custom background when it matters.
167
- 3. **Thresholding before GSEA** — GSEA needs the *full* ranked list; only ORA uses a cut list.
168
- 4. **Ranking GSEA by log2FoldChange alone** — unstable for low-count genes; prefer `stat` or `sign(LFC) * -log10(p)`.
169
- 5. **Multiple-testing across libraries** — FDR is computed *within* a library; running many libraries multiplies tests. Report per-library FDR and stay conservative.
170
- 6. **Redundant GO terms** — don't report 40 variants of the same term; collapse and show representatives.
171
- 7. **Significance ≠ relevance** — check the overlap count and gene-set size; tiny sets reach significance trivially.
172
- 8. **List too short/long for ORA** — <10 genes is underpowered; >2000 loses specificity (consider GSEA instead).
173
- 9. **No reproducibility metadata** — Enrichr/GO libraries are versioned and drift over time. Record library names+date and set a GSEA `seed`.
174
-
175
- ## Integration with Other Skills
176
-
177
- - **Upstream (where genes come from):** `pydeseq2` (DE genes + `stat` for GSEA), `scanpy` (`rank_genes_groups` markers / scores), `depmap`/`pytdc` (screen hits), proteomics skills (`pyopenms`, `matchms`).
178
- - **Databases / IDs:** `database-lookup` (Reactome, KEGG, STRING, Gene Ontology APIs), `gget` (`gget enrichr` quick path, `gget info` for ID mapping), `bioservices`.
179
- - **Downstream:** `scientific-visualization` (custom figures), `networkx` (enrichment-map graphs), `scientific-writing` / `literature-review` (interpret and cite), `statistical-analysis` (multiple-testing details).
180
-
181
- ## Reference Files
182
-
183
- Read the relevant file when you need depth:
184
-
185
- - `references/gseapy.md` — full gseapy API: `enrichr`, offline `enrich`, `prerank`, `gsea`, `ssgsea`, `gsva`, `Msigdb`, `Biomart`, `get_library_name`/`read_gmt`, every plot, result-column meanings, GMT/offline usage, and troubleshooting (rate limits, empty results).
186
- - `references/databases-and-gene-sets.md` — GO, KEGG, Reactome, WikiPathways, MSigDB collections, Enrichr library naming, g:Profiler sources, organism handling, gene-ID conversion, library selection by question, and pointers to Reactome/STRING APIs and decoupler activity inference.
187
- - `references/interpretation.md` — ORA vs GSEA statistics, background-universe choice, multiple-testing methods (BH vs g:SCS vs Bonferroni), leading-edge genes, redundancy reduction, effect vs significance, a publication-table template, and reproducibility checklist.
188
-
189
- ## Resources
190
-
191
- - gseapy docs: https://gseapy.readthedocs.io/ · repo: https://github.com/zqfang/GSEApy
192
- - g:Profiler: https://biit.cs.ut.ee/gprofiler/ · Python client: https://pypi.org/project/gprofiler-official/
193
- - Enrichr: https://maayanlab.cloud/Enrichr/ · MSigDB: https://www.gsea-msigdb.org/gsea/msigdb/
194
- - GSEA method: Subramanian et al. (2005) PNAS, DOI: 10.1073/pnas.0506580102
@@ -1,322 +0,0 @@
1
- ---
2
- name: pdf
3
- description: Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one, splitting PDFs apart, rotating pages, adding watermarks, creating new PDFs, filling PDF forms, encrypting/decrypting PDFs, extracting images, and OCR on scanned PDFs to make them searchable. If the user mentions a .pdf file or asks to produce one, use this skill.
4
- license: Proprietary. LICENSE.txt has complete terms
5
- metadata:
6
- version: "1.2"
7
- skill-author: Anthropic, PBC
8
- source: https://github.com/anthropics/skills/tree/main/skills/pdf
9
- ---
10
-
11
- # PDF Processing Guide
12
-
13
- ## Overview
14
-
15
- This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read forms.md and follow its instructions.
16
-
17
- ## Quick Start
18
-
19
- ```python
20
- from pypdf import PdfReader, PdfWriter
21
-
22
- # Read a PDF
23
- reader = PdfReader("document.pdf")
24
- print(f"Pages: {len(reader.pages)}")
25
-
26
- # Extract text
27
- text = ""
28
- for page in reader.pages:
29
- text += page.extract_text()
30
- ```
31
-
32
- ## Python Libraries
33
-
34
- ### pypdf - Basic Operations
35
-
36
- #### Merge PDFs
37
- ```python
38
- from pypdf import PdfWriter, PdfReader
39
-
40
- writer = PdfWriter()
41
- for pdf_file in ["doc1.pdf", "doc2.pdf", "doc3.pdf"]:
42
- reader = PdfReader(pdf_file)
43
- for page in reader.pages:
44
- writer.add_page(page)
45
-
46
- with open("merged.pdf", "wb") as output:
47
- writer.write(output)
48
- ```
49
-
50
- #### Split PDF
51
- ```python
52
- reader = PdfReader("input.pdf")
53
- for i, page in enumerate(reader.pages):
54
- writer = PdfWriter()
55
- writer.add_page(page)
56
- with open(f"page_{i+1}.pdf", "wb") as output:
57
- writer.write(output)
58
- ```
59
-
60
- #### Extract Metadata
61
- ```python
62
- reader = PdfReader("document.pdf")
63
- meta = reader.metadata
64
- print(f"Title: {meta.title}")
65
- print(f"Author: {meta.author}")
66
- print(f"Subject: {meta.subject}")
67
- print(f"Creator: {meta.creator}")
68
- ```
69
-
70
- #### Rotate Pages
71
- ```python
72
- reader = PdfReader("input.pdf")
73
- writer = PdfWriter()
74
-
75
- page = reader.pages[0]
76
- page.rotate(90) # Rotate 90 degrees clockwise
77
- writer.add_page(page)
78
-
79
- with open("rotated.pdf", "wb") as output:
80
- writer.write(output)
81
- ```
82
-
83
- ### pdfplumber - Text and Table Extraction
84
-
85
- #### Extract Text with Layout
86
- ```python
87
- import pdfplumber
88
-
89
- with pdfplumber.open("document.pdf") as pdf:
90
- for page in pdf.pages:
91
- text = page.extract_text()
92
- print(text)
93
- ```
94
-
95
- #### Extract Tables
96
- ```python
97
- with pdfplumber.open("document.pdf") as pdf:
98
- for i, page in enumerate(pdf.pages):
99
- tables = page.extract_tables()
100
- for j, table in enumerate(tables):
101
- print(f"Table {j+1} on page {i+1}:")
102
- for row in table:
103
- print(row)
104
- ```
105
-
106
- #### Advanced Table Extraction
107
- ```python
108
- import pandas as pd
109
-
110
- with pdfplumber.open("document.pdf") as pdf:
111
- all_tables = []
112
- for page in pdf.pages:
113
- tables = page.extract_tables()
114
- for table in tables:
115
- if table: # Check if table is not empty
116
- df = pd.DataFrame(table[1:], columns=table[0])
117
- all_tables.append(df)
118
-
119
- # Combine all tables
120
- if all_tables:
121
- combined_df = pd.concat(all_tables, ignore_index=True)
122
- combined_df.to_excel("extracted_tables.xlsx", index=False)
123
- ```
124
-
125
- ### reportlab - Create PDFs
126
-
127
- #### Basic PDF Creation
128
- ```python
129
- from reportlab.lib.pagesizes import letter
130
- from reportlab.pdfgen import canvas
131
-
132
- c = canvas.Canvas("hello.pdf", pagesize=letter)
133
- width, height = letter
134
-
135
- # Add text
136
- c.drawString(100, height - 100, "Hello World!")
137
- c.drawString(100, height - 120, "This is a PDF created with reportlab")
138
-
139
- # Add a line
140
- c.line(100, height - 140, 400, height - 140)
141
-
142
- # Save
143
- c.save()
144
- ```
145
-
146
- #### Create PDF with Multiple Pages
147
- ```python
148
- from reportlab.lib.pagesizes import letter
149
- from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, PageBreak
150
- from reportlab.lib.styles import getSampleStyleSheet
151
-
152
- doc = SimpleDocTemplate("report.pdf", pagesize=letter)
153
- styles = getSampleStyleSheet()
154
- story = []
155
-
156
- # Add content
157
- title = Paragraph("Report Title", styles['Title'])
158
- story.append(title)
159
- story.append(Spacer(1, 12))
160
-
161
- body = Paragraph("This is the body of the report. " * 20, styles['Normal'])
162
- story.append(body)
163
- story.append(PageBreak())
164
-
165
- # Page 2
166
- story.append(Paragraph("Page 2", styles['Heading1']))
167
- story.append(Paragraph("Content for page 2", styles['Normal']))
168
-
169
- # Build PDF
170
- doc.build(story)
171
- ```
172
-
173
- #### Subscripts and Superscripts
174
-
175
- **IMPORTANT**: Never use Unicode subscript/superscript characters (₀₁₂₃₄₅₆₇₈₉, ⁰¹²³⁴⁵⁶⁷⁸⁹) in ReportLab PDFs. The built-in fonts do not include these glyphs, causing them to render as solid black boxes.
176
-
177
- Instead, use ReportLab's XML markup tags in Paragraph objects:
178
- ```python
179
- from reportlab.platypus import Paragraph
180
- from reportlab.lib.styles import getSampleStyleSheet
181
-
182
- styles = getSampleStyleSheet()
183
-
184
- # Subscripts: use <sub> tag
185
- chemical = Paragraph("H<sub>2</sub>O", styles['Normal'])
186
-
187
- # Superscripts: use <super> tag
188
- squared = Paragraph("x<super>2</super> + y<super>2</super>", styles['Normal'])
189
- ```
190
-
191
- For canvas-drawn text (not Paragraph objects), manually adjust font the size and position rather than using Unicode subscripts/superscripts.
192
-
193
- ## Command-Line Tools
194
-
195
- ### pdftotext (poppler-utils)
196
- ```bash
197
- # Extract text
198
- pdftotext input.pdf output.txt
199
-
200
- # Extract text preserving layout
201
- pdftotext -layout input.pdf output.txt
202
-
203
- # Extract specific pages
204
- pdftotext -f 1 -l 5 input.pdf output.txt # Pages 1-5
205
- ```
206
-
207
- ### qpdf
208
- ```bash
209
- # Merge PDFs
210
- qpdf --empty --pages file1.pdf file2.pdf -- merged.pdf
211
-
212
- # Split pages
213
- qpdf input.pdf --pages . 1-5 -- pages1-5.pdf
214
- qpdf input.pdf --pages . 6-10 -- pages6-10.pdf
215
-
216
- # Rotate pages
217
- qpdf input.pdf output.pdf --rotate=+90:1 # Rotate page 1 by 90 degrees
218
-
219
- # Remove password
220
- qpdf --password=mypassword --decrypt encrypted.pdf decrypted.pdf
221
- ```
222
-
223
- ### pdftk (if available)
224
- ```bash
225
- # Merge
226
- pdftk file1.pdf file2.pdf cat output merged.pdf
227
-
228
- # Split
229
- pdftk input.pdf burst
230
-
231
- # Rotate
232
- pdftk input.pdf rotate 1east output rotated.pdf
233
- ```
234
-
235
- ## Common Tasks
236
-
237
- ### Extract Text from Scanned PDFs
238
- ```python
239
- # Requires: uv pip install pytesseract pdf2image
240
- import pytesseract
241
- from pdf2image import convert_from_path
242
-
243
- # Convert PDF to images
244
- images = convert_from_path('scanned.pdf')
245
-
246
- # OCR each page
247
- text = ""
248
- for i, image in enumerate(images):
249
- text += f"Page {i+1}:\n"
250
- text += pytesseract.image_to_string(image)
251
- text += "\n\n"
252
-
253
- print(text)
254
- ```
255
-
256
- ### Add Watermark
257
- ```python
258
- from pypdf import PdfReader, PdfWriter
259
-
260
- # Create watermark (or load existing)
261
- watermark = PdfReader("watermark.pdf").pages[0]
262
-
263
- # Apply to all pages
264
- reader = PdfReader("document.pdf")
265
- writer = PdfWriter()
266
-
267
- for page in reader.pages:
268
- page.merge_page(watermark)
269
- writer.add_page(page)
270
-
271
- with open("watermarked.pdf", "wb") as output:
272
- writer.write(output)
273
- ```
274
-
275
- ### Extract Images
276
- ```bash
277
- # Using pdfimages (poppler-utils)
278
- pdfimages -j input.pdf output_prefix
279
-
280
- # This extracts all images as output_prefix-000.jpg, output_prefix-001.jpg, etc.
281
- ```
282
-
283
- ### Password Protection
284
- ```python
285
- from pypdf import PdfReader, PdfWriter
286
-
287
- reader = PdfReader("input.pdf")
288
- writer = PdfWriter()
289
-
290
- for page in reader.pages:
291
- writer.add_page(page)
292
-
293
- # Add password
294
- writer.encrypt("userpassword", "ownerpassword")
295
-
296
- with open("encrypted.pdf", "wb") as output:
297
- writer.write(output)
298
- ```
299
-
300
- ## Quick Reference
301
-
302
- | Task | Best Tool | Command/Code |
303
- |------|-----------|--------------|
304
- | Merge PDFs | pypdf | `writer.add_page(page)` |
305
- | Split PDFs | pypdf | One page per file |
306
- | Extract text | pdfplumber | `page.extract_text()` |
307
- | Extract tables | pdfplumber | `page.extract_tables()` |
308
- | Create PDFs | reportlab | Canvas or Platypus |
309
- | Command line merge | qpdf | `qpdf --empty --pages ...` |
310
- | OCR scanned PDFs | pytesseract | Convert to image first |
311
- | Fill PDF forms | pdf-lib or pypdf (see forms.md) | See forms.md |
312
-
313
- ## Next Steps
314
-
315
- - For advanced pypdfium2 usage, see reference.md
316
- - For JavaScript libraries (pdf-lib), see reference.md
317
- - If you need to fill out a PDF form, follow the instructions in forms.md
318
- - For troubleshooting guides, see reference.md
319
-
320
- ---
321
-
322
- *This skill is created and maintained by [Anthropic](https://github.com/anthropics/skills/tree/main/skills/pdf). Vendored here unmodified except for frontmatter metadata and the case of the `reference.md`/`forms.md` links, which upstream writes uppercase; see LICENSE.txt for terms.*