@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,369 +0,0 @@
1
- ---
2
- name: pydeseq2
3
- description: Differential gene expression analysis for bulk RNA-seq with PyDESeq2, including formulaic designs, Wald tests, FDR correction, LFC shrinkage, and result visualization.
4
- allowed-tools: Read Write Edit Bash
5
- compatibility: Requires Python >=3.11 and PyDESeq2 0.5.4-compatible dependencies. Examples target PyDESeq2 0.5.x, formulaic design strings, explicit contrasts, and uv-based installs.
6
- license: MIT license
7
- metadata:
8
- version: "1.3"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # PyDESeq2
13
-
14
- ## Overview
15
-
16
- PyDESeq2 is a Python implementation of DESeq2 for differential expression analysis with bulk RNA-seq data. Design and execute complete workflows from data loading through result interpretation, including formulaic single-factor and multi-factor designs, Wald tests with multiple testing correction, optional apeGLM shrinkage, and integration with pandas and AnnData.
17
-
18
- ## When to Use This Skill
19
-
20
- This skill should be used when:
21
- - Analyzing bulk RNA-seq count data for differential expression
22
- - Comparing gene expression between experimental conditions (e.g., treated vs control)
23
- - Performing multi-factor designs accounting for batch effects or covariates
24
- - Converting R-based DESeq2 workflows to Python
25
- - Integrating differential expression analysis into Python-based pipelines
26
- - Users mention "DESeq2", "differential expression", "RNA-seq analysis", or "PyDESeq2"
27
-
28
- ## Quick Start Workflow
29
-
30
- For users who want to perform a standard differential expression analysis:
31
-
32
- ```python
33
- import pandas as pd
34
- from pydeseq2.dds import DeseqDataSet
35
- from pydeseq2.default_inference import DefaultInference
36
- from pydeseq2.ds import DeseqStats
37
-
38
- # 1. Load data
39
- counts_df = pd.read_csv("counts.csv", index_col=0).T # Transpose to samples × genes
40
- metadata = pd.read_csv("metadata.csv", index_col=0)
41
-
42
- # 2. Filter low-count genes
43
- genes_to_keep = counts_df.columns[counts_df.sum(axis=0) >= 10]
44
- counts_df = counts_df[genes_to_keep]
45
-
46
- # 3. Make the reference level explicit and fit DESeq2
47
- metadata["condition"] = pd.Categorical(
48
- metadata["condition"], categories=["control", "treated"]
49
- )
50
- inference = DefaultInference(n_cpus=4)
51
- dds = DeseqDataSet(
52
- counts=counts_df,
53
- metadata=metadata,
54
- design="~condition",
55
- refit_cooks=True,
56
- inference=inference,
57
- )
58
- dds.deseq2()
59
-
60
- # 4. Perform statistical testing
61
- ds = DeseqStats(
62
- dds,
63
- contrast=["condition", "treated", "control"],
64
- inference=inference,
65
- )
66
- ds.summary()
67
-
68
- # 5. Access results
69
- results = ds.results_df
70
- significant = results[results.padj < 0.05]
71
- print(f"Found {len(significant)} significant genes")
72
- ```
73
-
74
- ## Core Workflow Steps
75
-
76
- The six steps, with code, are in
77
- [references/core_workflow_steps.md](references/core_workflow_steps.md):
78
-
79
- 1. **Data preparation** — raw integer counts with genes as columns and samples as rows,
80
- and matching metadata. Never feed normalized or transformed values to DESeq2.
81
- 2. **Design specification** — the design factors and the reference level for each.
82
- 3. **DESeq2 fitting** — size factors, dispersions, and the GLM fit.
83
- 4. **Statistical testing** — Wald tests for a named contrast.
84
- 5. **Optional LFC shrinkage** — for ranking and visualization.
85
- 6. **Result export** — the results table with adjusted p-values.
86
-
87
- Multi-factor designs, contrasts, and interaction terms are in
88
- [references/analysis_patterns.md](references/analysis_patterns.md).
89
-
90
- ## Using the Analysis Script
91
-
92
- This skill includes a complete command-line script for standard analyses:
93
-
94
- ```bash
95
- # Basic usage
96
- python scripts/run_deseq2_analysis.py \
97
- --counts counts.csv \
98
- --metadata metadata.csv \
99
- --design "~condition" \
100
- --contrast condition treated control \
101
- --output results/
102
-
103
- # With additional options
104
- python scripts/run_deseq2_analysis.py \
105
- --counts counts.csv \
106
- --metadata metadata.csv \
107
- --design "~batch + condition" \
108
- --contrast condition treated control \
109
- --output results/ \
110
- --min-counts 10 \
111
- --alpha 0.05 \
112
- --n-cpus 4 \
113
- --shrink-coeff "condition[T.treated]" \
114
- --plots
115
- ```
116
-
117
- **Script features:**
118
- - Automatic data loading and validation
119
- - Gene and sample filtering
120
- - Complete DESeq2 pipeline execution
121
- - Statistical testing with customizable parameters
122
- - Result export (CSV and portable AnnData/H5AD)
123
- - Explicit LFC shrinkage coefficient support for PyDESeq2 0.5.x
124
- - Optional visualization (volcano and MA plots)
125
-
126
- Refer users to `scripts/run_deseq2_analysis.py` when they need a standalone analysis tool or want to batch process multiple datasets.
127
-
128
- ## Result Interpretation
129
-
130
- ### Identifying Significant Genes
131
-
132
- ```python
133
- # Filter by adjusted p-value
134
- significant = ds.results_df[ds.results_df.padj < 0.05]
135
-
136
- # Filter by both significance and effect size
137
- sig_and_large = ds.results_df[
138
- (ds.results_df.padj < 0.05) &
139
- (abs(ds.results_df.log2FoldChange) > 1)
140
- ]
141
-
142
- # Separate up- and down-regulated
143
- upregulated = significant[significant.log2FoldChange > 0]
144
- downregulated = significant[significant.log2FoldChange < 0]
145
-
146
- print(f"Upregulated: {len(upregulated)}")
147
- print(f"Downregulated: {len(downregulated)}")
148
- ```
149
-
150
- ### Ranking and Sorting
151
-
152
- ```python
153
- # Sort by adjusted p-value
154
- top_by_padj = ds.results_df.sort_values("padj").head(20)
155
-
156
- # Sort by absolute fold change (use shrunk values)
157
- ds.lfc_shrink(coeff="condition[T.treated]")
158
- ds.results_df["abs_lfc"] = abs(ds.results_df.log2FoldChange)
159
- top_by_lfc = ds.results_df.sort_values("abs_lfc", ascending=False).head(20)
160
-
161
- # Sort by a combined metric
162
- ds.results_df["score"] = -np.log10(ds.results_df.padj) * abs(ds.results_df.log2FoldChange)
163
- top_combined = ds.results_df.sort_values("score", ascending=False).head(20)
164
- ```
165
-
166
- ### Quality Metrics
167
-
168
- ```python
169
- # Check normalization (size factors should be close to 1)
170
- print("Size factors:", dds.obs["size_factors"])
171
-
172
- # Examine dispersion estimates
173
- import matplotlib.pyplot as plt
174
- plt.hist(dds.var["dispersions"], bins=50)
175
- plt.xlabel("Dispersion")
176
- plt.ylabel("Frequency")
177
- plt.title("Dispersion Distribution")
178
- plt.show()
179
-
180
- # Check p-value distribution (should be mostly flat with peak near 0)
181
- plt.hist(ds.results_df.pvalue.dropna(), bins=50)
182
- plt.xlabel("P-value")
183
- plt.ylabel("Frequency")
184
- plt.title("P-value Distribution")
185
- plt.show()
186
- ```
187
-
188
- ## Visualization Guidelines
189
-
190
- ### Volcano Plot
191
-
192
- Visualize significance vs effect size:
193
-
194
- ```python
195
- import matplotlib.pyplot as plt
196
- import numpy as np
197
-
198
- results = ds.results_df.copy()
199
- results["-log10(padj)"] = -np.log10(results.padj)
200
-
201
- plt.figure(figsize=(10, 6))
202
- significant = results.padj < 0.05
203
-
204
- plt.scatter(
205
- results.loc[~significant, "log2FoldChange"],
206
- results.loc[~significant, "-log10(padj)"],
207
- alpha=0.3, s=10, c='gray', label='Not significant'
208
- )
209
- plt.scatter(
210
- results.loc[significant, "log2FoldChange"],
211
- results.loc[significant, "-log10(padj)"],
212
- alpha=0.6, s=10, c='red', label='padj < 0.05'
213
- )
214
-
215
- plt.axhline(-np.log10(0.05), color='blue', linestyle='--', alpha=0.5)
216
- plt.xlabel("Log2 Fold Change")
217
- plt.ylabel("-Log10(Adjusted P-value)")
218
- plt.title("Volcano Plot")
219
- plt.legend()
220
- plt.savefig("volcano_plot.png", dpi=300)
221
- ```
222
-
223
- ### MA Plot
224
-
225
- Show fold change vs mean expression:
226
-
227
- ```python
228
- plt.figure(figsize=(10, 6))
229
-
230
- plt.scatter(
231
- np.log10(results.loc[~significant, "baseMean"] + 1),
232
- results.loc[~significant, "log2FoldChange"],
233
- alpha=0.3, s=10, c='gray'
234
- )
235
- plt.scatter(
236
- np.log10(results.loc[significant, "baseMean"] + 1),
237
- results.loc[significant, "log2FoldChange"],
238
- alpha=0.6, s=10, c='red'
239
- )
240
-
241
- plt.axhline(0, color='blue', linestyle='--', alpha=0.5)
242
- plt.xlabel("Log10(Base Mean + 1)")
243
- plt.ylabel("Log2 Fold Change")
244
- plt.title("MA Plot")
245
- plt.savefig("ma_plot.png", dpi=300)
246
- ```
247
-
248
- ## Troubleshooting Common Issues
249
-
250
- ### Data Format Problems
251
-
252
- **Issue:** "Index mismatch between counts and metadata"
253
-
254
- **Solution:** Ensure sample names match exactly
255
- ```python
256
- print("Counts samples:", counts_df.index.tolist())
257
- print("Metadata samples:", metadata.index.tolist())
258
-
259
- # Take intersection if needed
260
- common = counts_df.index.intersection(metadata.index)
261
- counts_df = counts_df.loc[common]
262
- metadata = metadata.loc[common]
263
- ```
264
-
265
- **Issue:** "All genes have zero counts"
266
-
267
- **Solution:** Check if data needs transposition
268
- ```python
269
- print(f"Counts shape: {counts_df.shape}")
270
- # If genes > samples, transpose is needed
271
- if counts_df.shape[1] < counts_df.shape[0]:
272
- counts_df = counts_df.T
273
- ```
274
-
275
- ### Design Matrix Issues
276
-
277
- **Issue:** "Design matrix is not full rank"
278
-
279
- **Cause:** Confounded variables (e.g., all treated samples in one batch)
280
-
281
- **Solution:** Remove confounded variable or add interaction term
282
- ```python
283
- # Check confounding
284
- print(pd.crosstab(metadata.condition, metadata.batch))
285
-
286
- # Either simplify design or add interaction
287
- design = "~condition" # Remove batch
288
- # OR
289
- design = "~condition + batch + condition:batch" # Model interaction
290
- ```
291
-
292
- ### No Significant Genes
293
-
294
- **Diagnostics:**
295
- ```python
296
- # Check dispersion distribution
297
- plt.hist(dds.var["dispersions"], bins=50)
298
- plt.show()
299
-
300
- # Check size factors
301
- print(dds.obs["size_factors"])
302
-
303
- # Look at top genes by raw p-value
304
- print(ds.results_df.nsmallest(20, "pvalue"))
305
- ```
306
-
307
- **Possible causes:**
308
- - Small effect sizes
309
- - High biological variability
310
- - Insufficient sample size
311
- - Technical issues (batch effects, outliers)
312
-
313
- ## Reference Documentation
314
-
315
- For comprehensive details beyond this workflow-oriented guide:
316
-
317
- - **API Reference** (`references/api_reference.md`): Complete documentation of PyDESeq2 classes, methods, and data structures. Use when needing detailed parameter information or understanding object attributes.
318
-
319
- - **Workflow Guide** (`references/workflow_guide.md`): In-depth guide covering complete analysis workflows, data loading patterns, multi-factor designs, troubleshooting, and best practices. Use when handling complex experimental designs or encountering issues.
320
-
321
- Load these references into context when users need:
322
- - Detailed API documentation: `Read references/api_reference.md`
323
- - Comprehensive workflow examples: `Read references/workflow_guide.md`
324
- - Troubleshooting guidance: `Read references/workflow_guide.md` (see Troubleshooting section)
325
-
326
- ## Key Reminders
327
-
328
- 1. **Data orientation matters:** Count matrices typically load as genes × samples but need to be samples × genes. Always transpose with `.T` if needed.
329
-
330
- 2. **Sample filtering:** Remove samples with missing metadata before analysis to avoid errors.
331
-
332
- 3. **Gene filtering:** Filter low-count genes (e.g., < 10 total reads) to improve power and reduce computational time.
333
-
334
- 4. **Design formula order:** Put adjustment variables before the variable of interest (e.g., `"~batch + condition"` not `"~condition + batch"`).
335
-
336
- 5. **LFC shrinkage timing:** Apply shrinkage after statistical testing and only for visualization/ranking purposes. P-values remain based on unshrunken estimates.
337
-
338
- 6. **Result interpretation:** Use `padj < 0.05` for significance, not raw p-values. The Benjamini-Hochberg procedure controls false discovery rate.
339
-
340
- 7. **Contrast specification:** The format is `[variable, test_level, reference_level]` where test_level is compared against reference_level.
341
-
342
- 8. **Save intermediate objects:** Prefer `dds.to_picklable_anndata().write_h5ad("dds_result.h5ad")` for portable outputs. Only load pickle files that you created yourself and trust.
343
-
344
- ## Installation and Requirements
345
-
346
- ```bash
347
- uv pip install pydeseq2==0.5.4
348
- ```
349
-
350
- **System requirements:**
351
- - Python 3.11+
352
- - PyDESeq2 0.5.4
353
- - pandas 2.2.0+
354
- - numpy 2.0.0+
355
- - scipy 1.12.0+
356
- - scikit-learn 1.4.0+
357
- - anndata 0.11.0+
358
- - formulaic 1.0.2+ and formulaic-contrasts 0.2.0+
359
-
360
- **Optional for visualization:**
361
- - matplotlib
362
- - seaborn
363
-
364
- ## Additional Resources
365
-
366
- - **Official Documentation:** https://pydeseq2.readthedocs.io
367
- - **GitHub Repository:** https://github.com/scverse/PyDESeq2
368
- - **Publication:** Muzellec et al. (2023) Bioinformatics, DOI: 10.1093/bioinformatics/btad547
369
- - **Original DESeq2 (R):** Love et al. (2014) Genome Biology, DOI: 10.1186/s13059-014-0550-8