@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,412 +0,0 @@
1
- ---
2
- name: deeptools
3
- description: NGS analysis toolkit. BAM to bigWig conversion, QC (correlation, PCA, fingerprints), heatmaps/profiles (TSS, peaks), for ChIP-seq, RNA-seq, ATAC-seq visualization.
4
- license: BSD license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python >3.8 and deepTools 3.5.6-compatible dependencies. The upstream project recommends conda/bioconda for full dependency resolution; repo examples use uv with pinned PyPI installs for reproducible command-line workflows.
7
- metadata:
8
- version: "1.2"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # deepTools: NGS Data Analysis Toolkit
13
-
14
- ## Overview
15
-
16
- deepTools is a comprehensive suite of Python command-line tools designed for processing and analyzing high-throughput sequencing data. Use deepTools to perform quality control, normalize data, compare samples, and generate publication-quality visualizations for ChIP-seq, RNA-seq, ATAC-seq, MNase-seq, and other NGS experiments.
17
-
18
- **Core capabilities:**
19
- - Convert BAM alignments to normalized coverage tracks (bigWig/bedGraph)
20
- - Quality control assessment (fingerprint, correlation, coverage)
21
- - Sample comparison and correlation analysis
22
- - Heatmap and profile plot generation around genomic features
23
- - Enrichment analysis and peak region visualization
24
-
25
- ## When to Use This Skill
26
-
27
- This skill should be used when:
28
-
29
- - **File conversion**: "Convert BAM to bigWig", "generate coverage tracks", "normalize ChIP-seq data"
30
- - **Quality control**: "check ChIP quality", "compare replicates", "assess sequencing depth", "QC analysis"
31
- - **Visualization**: "create heatmap around TSS", "plot ChIP signal", "visualize enrichment", "generate profile plot"
32
- - **Sample comparison**: "compare treatment vs control", "correlate samples", "PCA analysis"
33
- - **Analysis workflows**: "analyze ChIP-seq data", "RNA-seq coverage", "ATAC-seq analysis", "complete workflow"
34
- - **Working with specific file types**: BAM files, bigWig files, BED region files in genomics context
35
-
36
- ## Quick Start
37
-
38
- For users new to deepTools, start with file validation and common workflows:
39
-
40
- ### 1. Validate Input Files
41
-
42
- Before running any analysis, validate BAM, bigWig, and BED files using the validation script:
43
-
44
- ```bash
45
- python scripts/validate_files.py --bam sample1.bam sample2.bam --bed regions.bed
46
- ```
47
-
48
- This checks file existence, BAM indices, and format correctness.
49
-
50
- ### 2. Generate Workflow Template
51
-
52
- For standard analyses, use the workflow generator to create customized scripts:
53
-
54
- ```bash
55
- # List available workflows
56
- python scripts/workflow_generator.py --list
57
-
58
- # Generate ChIP-seq QC workflow
59
- python scripts/workflow_generator.py chipseq_qc -o qc_workflow.sh \
60
- --input-bam Input.bam --chip-bams "ChIP1.bam ChIP2.bam" \
61
- --genome-size 2913022398
62
-
63
- # Make executable and run
64
- chmod +x qc_workflow.sh
65
- ./qc_workflow.sh
66
- ```
67
-
68
- ### 3. Most Common Operations
69
-
70
- See `assets/quick_reference.md` for frequently used commands and parameters.
71
-
72
- ## Installation
73
-
74
- ```bash
75
- uv pip install deepTools==3.5.6
76
- ```
77
-
78
- Upstream recommends conda/bioconda for full dependency resolution, especially on shared HPC systems:
79
-
80
- ```bash
81
- conda install -c conda-forge -c bioconda deeptools
82
- ```
83
-
84
- On Apple Silicon, upstream documents either the PyPI route above or an `osx-64` conda environment when native conda packages are unavailable.
85
-
86
- ## Core Workflows and Tool Categories
87
-
88
- Complete command sequences for ChIP-seq QC, full ChIP-seq analysis, RNA-seq coverage, and
89
- ATAC-seq analysis — plus the BAM/bigWig processing, quality control, and visualization
90
- tool categories — are in [references/core_workflows.md](references/core_workflows.md) and
91
- [references/workflows.md](references/workflows.md). Per-tool options are in
92
- [references/tools_reference.md](references/tools_reference.md).
93
-
94
- ## Normalization Methods
95
-
96
- Choosing the correct normalization is critical for valid comparisons. Consult `references/normalization_methods.md` for comprehensive guidance.
97
-
98
- **Quick selection guide:**
99
-
100
- - **ChIP-seq coverage**: Use RPGC or CPM
101
- - **ChIP-seq comparison**: Use bamCompare with log2 and readCount
102
- - **RNA-seq bins**: Use CPM
103
- - **RNA-seq genes**: Use RPKM (accounts for gene length)
104
- - **ATAC-seq**: Use RPGC or CPM
105
-
106
- **Normalization methods:**
107
- - **RPGC**: 1× genome coverage (requires --effectiveGenomeSize)
108
- - **CPM**: Counts per million mapped reads
109
- - **RPKM**: Reads per kb per million (per-bin length and library-size scaling)
110
- - **BPM**: Bins per million, analogous to TPM-style scaling over binned signal
111
- - **None**: Raw counts (not recommended for comparisons)
112
-
113
- Full explanation: `references/normalization_methods.md`
114
-
115
- ## Effective Genome Sizes
116
-
117
- RPGC normalization requires effective genome size. Common values:
118
-
119
- | Organism | Assembly | Size | Usage |
120
- |----------|----------|------|-------|
121
- | Human | GRCh38/hg38 | 2,913,022,398 | `--effectiveGenomeSize 2913022398` |
122
- | Human | T2T/CHM13CAT_v2 | 3,117,292,070 | `--effectiveGenomeSize 3117292070` |
123
- | Mouse | GRCm39/mm39 | 2,654,621,783 | `--effectiveGenomeSize 2654621783` |
124
- | Mouse | GRCm38/mm10 | 2,652,783,500 | `--effectiveGenomeSize 2652783500` |
125
- | Zebrafish | GRCz11 | 1,368,780,147 | `--effectiveGenomeSize 1368780147` |
126
- | *Drosophila* | dm6 | 142,573,017 | `--effectiveGenomeSize 142573017` |
127
- | *C. elegans* | ce10/ce11 | 100,286,401 | `--effectiveGenomeSize 100286401` |
128
-
129
- Complete table with read-length-specific values: `references/effective_genome_sizes.md`
130
-
131
- ## Common Parameters Across Tools
132
-
133
- Many deepTools commands share these options:
134
-
135
- **Performance:**
136
- - `--numberOfProcessors, -p`: Enable parallel processing (always use available cores)
137
- - `max` / `max/2`: Supported values for `--numberOfProcessors`; useful under schedulers because recent deepTools releases detect CPU affinity more carefully
138
- - `--region`: Process specific regions for testing (e.g., `chr1:1-1000000`)
139
-
140
- **Read Filtering:**
141
- - `--ignoreDuplicates`: Remove PCR duplicates (recommended for most analyses)
142
- - `--minMappingQuality`: Filter by alignment quality (e.g., `--minMappingQuality 10`)
143
- - `--minFragmentLength` / `--maxFragmentLength`: Fragment length bounds
144
- - `--samFlagInclude` / `--samFlagExclude`: SAM flag filtering
145
-
146
- **Read Processing:**
147
- - `--extendReads`: Extend to fragment length (ChIP-seq: YES, RNA-seq: NO)
148
- - `--centerReads`: Center at fragment midpoint for sharper signals
149
-
150
- ## Best Practices
151
-
152
- ### File Validation
153
- **Always validate files first** using `scripts/validate_files.py` to check:
154
- - File existence and readability
155
- - BAM indices present (.bai files)
156
- - BED format correctness
157
- - File sizes reasonable
158
-
159
- ### Analysis Strategy
160
-
161
- 1. **Start with QC**: Run correlation, coverage, and fingerprint analysis before proceeding
162
- 2. **Test on small regions**: Use `--region chr1:1-10000000` for parameter testing
163
- 3. **Document commands**: Save full command lines for reproducibility
164
- 4. **Use consistent normalization**: Apply same method across samples in comparisons
165
- 5. **Verify genome assembly**: Ensure BAM and BED files use matching genome builds
166
-
167
- ### ChIP-seq Specific
168
-
169
- - **Always extend reads** for ChIP-seq: `--extendReads 200`
170
- - **Remove duplicates**: Use `--ignoreDuplicates` in most cases
171
- - **Check enrichment first**: Run plotFingerprint before detailed analysis
172
- - **GC correction**: Only apply if significant bias detected; never use `--ignoreDuplicates` after GC correction
173
-
174
- ### RNA-seq Specific
175
-
176
- - **Never extend reads** for RNA-seq (would span splice junctions)
177
- - **Strand-specific**: Use `--filterRNAstrand forward/reverse` for common dUTP-style stranded libraries; confirm library orientation before interpreting strand labels
178
- - **Normalization**: CPM for bins, RPKM for genes
179
-
180
- ### ATAC-seq Specific
181
-
182
- - **Apply Tn5 correction**: Use alignmentSieve with `--ATACshift`
183
- - **Use only proper pairs for shifting**: `--ATACshift` is equivalent to `--shift 4 -5 5 -4` and filters to properly paired fragments
184
- - **Fragment filtering**: Set appropriate min/max fragment lengths
185
- - **Check nucleosome pattern**: Fragment size plot should show ladder pattern
186
-
187
- ### Performance Optimization
188
-
189
- 1. **Use multiple processors**: `--numberOfProcessors 8` (or available cores)
190
- 2. **Increase bin size** for faster processing and smaller files
191
- 3. **Process chromosomes separately** for memory-limited systems
192
- 4. **Pre-filter BAM files** using alignmentSieve to create reusable filtered files
193
- 5. **Use bigWig over bedGraph**: Compressed and faster to process
194
-
195
- ## Troubleshooting
196
-
197
- ### Common Issues
198
-
199
- **BAM index missing:**
200
- ```bash
201
- samtools index input.bam
202
- ```
203
-
204
- **Out of memory:**
205
- Process chromosomes individually using `--region`:
206
- ```bash
207
- bamCoverage --bam input.bam -o chr1.bw --region chr1
208
- ```
209
-
210
- **Slow processing:**
211
- Increase `--numberOfProcessors` and/or increase `--binSize`
212
-
213
- **bigWig files too large:**
214
- Increase bin size: `--binSize 50` or larger
215
-
216
- ### Validation Errors
217
-
218
- Run validation script to identify issues:
219
- ```bash
220
- python scripts/validate_files.py --bam *.bam --bed regions.bed
221
- ```
222
-
223
- Common errors and solutions explained in script output.
224
-
225
- ## Reference Documentation
226
-
227
- This skill includes comprehensive reference documentation:
228
-
229
- ### references/tools_reference.md
230
- Complete documentation of all deepTools commands organized by category:
231
- - BAM and bigWig processing tools (9 tools)
232
- - Quality control tools (6 tools)
233
- - Visualization tools (3 tools)
234
- - Miscellaneous tools (3 tools, including `bigwigAverage`)
235
-
236
- Each tool includes:
237
- - Purpose and overview
238
- - Key parameters with explanations
239
- - Usage examples
240
- - Important notes and best practices
241
-
242
- **Use this reference when:** Users ask about specific tools, parameters, or detailed usage.
243
-
244
- ### references/workflows.md
245
- Complete workflow examples for common analyses:
246
- - ChIP-seq quality control workflow
247
- - ChIP-seq complete analysis workflow
248
- - RNA-seq coverage workflow
249
- - ATAC-seq analysis workflow
250
- - Multi-sample comparison workflow
251
- - Peak region analysis workflow
252
- - Troubleshooting and performance tips
253
-
254
- **Use this reference when:** Users need complete analysis pipelines or workflow examples.
255
-
256
- ### references/normalization_methods.md
257
- Comprehensive guide to normalization methods:
258
- - Detailed explanation of each method (RPGC, CPM, RPKM, BPM, etc.)
259
- - When to use each method
260
- - Formulas and interpretation
261
- - Selection guide by experiment type
262
- - Common pitfalls and solutions
263
- - Quick reference table
264
-
265
- **Use this reference when:** Users ask about normalization, comparing samples, or which method to use.
266
-
267
- ### references/effective_genome_sizes.md
268
- Effective genome size values and usage:
269
- - Common organism values (human, mouse, fly, worm, zebrafish)
270
- - Read-length-specific values
271
- - Calculation methods
272
- - When and how to use in commands
273
- - Custom genome calculation instructions
274
-
275
- **Use this reference when:** Users need genome size for RPGC normalization or GC bias correction.
276
-
277
- ## Helper Scripts
278
-
279
- ### scripts/validate_files.py
280
-
281
- Validates BAM, bigWig, and BED files for deepTools analysis. Checks file existence, indices, and format.
282
-
283
- **Usage:**
284
- ```bash
285
- python scripts/validate_files.py --bam sample1.bam sample2.bam \
286
- --bed peaks.bed --bigwig signal.bw
287
- ```
288
-
289
- **When to use:** Before starting any analysis, or when troubleshooting errors.
290
-
291
- ### scripts/workflow_generator.py
292
-
293
- Generates customizable bash script templates for common deepTools workflows.
294
-
295
- **Available workflows:**
296
- - `chipseq_qc`: ChIP-seq quality control
297
- - `chipseq_analysis`: Complete ChIP-seq analysis
298
- - `rnaseq_coverage`: Strand-specific RNA-seq coverage
299
- - `atacseq`: ATAC-seq with Tn5 correction
300
-
301
- **Usage:**
302
- ```bash
303
- # List workflows
304
- python scripts/workflow_generator.py --list
305
-
306
- # Generate workflow
307
- python scripts/workflow_generator.py chipseq_qc -o qc.sh \
308
- --input-bam Input.bam --chip-bams "ChIP1.bam ChIP2.bam" \
309
- --genome-size 2913022398 --threads 8
310
-
311
- # Run generated workflow
312
- chmod +x qc.sh
313
- ./qc.sh
314
- ```
315
-
316
- **When to use:** Users request standard workflows or need template scripts to customize.
317
-
318
- ## Assets
319
-
320
- ### assets/quick_reference.md
321
-
322
- Quick reference card with most common commands, effective genome sizes, and typical workflow pattern.
323
-
324
- **When to use:** Users need quick command examples without detailed documentation.
325
-
326
- ## Handling User Requests
327
-
328
- ### For New Users
329
-
330
- 1. Start with installation verification
331
- 2. Validate input files using `scripts/validate_files.py`
332
- 3. Recommend appropriate workflow based on experiment type
333
- 4. Generate workflow template using `scripts/workflow_generator.py`
334
- 5. Guide through customization and execution
335
-
336
- ### For Experienced Users
337
-
338
- 1. Provide specific tool commands for requested operations
339
- 2. Reference appropriate sections in `references/tools_reference.md`
340
- 3. Suggest optimizations and best practices
341
- 4. Offer troubleshooting for issues
342
-
343
- ### For Specific Tasks
344
-
345
- **"Convert BAM to bigWig":**
346
- - Use bamCoverage with appropriate normalization
347
- - Recommend RPGC or CPM based on use case
348
- - Provide effective genome size for organism
349
- - Suggest relevant parameters (extendReads, ignoreDuplicates, binSize)
350
-
351
- **"Check ChIP quality":**
352
- - Run full QC workflow or use plotFingerprint specifically
353
- - Explain interpretation of results
354
- - Suggest follow-up actions based on results
355
-
356
- **"Create heatmap":**
357
- - Guide through two-step process: computeMatrix → plotHeatmap
358
- - Help choose appropriate matrix mode (reference-point vs scale-regions)
359
- - Suggest visualization parameters and clustering options
360
-
361
- **"Compare samples":**
362
- - Recommend bamCompare for two-sample comparison
363
- - Suggest multiBamSummary + plotCorrelation for multiple samples
364
- - Guide normalization method selection
365
-
366
- ### Referencing Documentation
367
-
368
- When users need detailed information:
369
- - **Tool details**: Direct to specific sections in `references/tools_reference.md`
370
- - **Workflows**: Use `references/workflows.md` for complete analysis pipelines
371
- - **Normalization**: Consult `references/normalization_methods.md` for method selection
372
- - **Genome sizes**: Reference `references/effective_genome_sizes.md`
373
-
374
- ## Example Interactions
375
-
376
- **User: "I need to analyze my ChIP-seq data"**
377
-
378
- Response approach:
379
- 1. Ask about files available (BAM files, peaks, genes)
380
- 2. Validate files using validation script
381
- 3. Generate chipseq_analysis workflow template
382
- 4. Customize for their specific files and organism
383
- 5. Explain each step as script runs
384
-
385
- **User: "Which normalization should I use?"**
386
-
387
- Response approach:
388
- 1. Ask about experiment type (ChIP-seq, RNA-seq, etc.)
389
- 2. Ask about comparison goal (within-sample or between-sample)
390
- 3. Consult `references/normalization_methods.md` selection guide
391
- 4. Recommend appropriate method with justification
392
- 5. Provide command example with parameters
393
-
394
- **User: "Create a heatmap around TSS"**
395
-
396
- Response approach:
397
- 1. Verify bigWig and gene BED files available
398
- 2. Use computeMatrix with reference-point mode at TSS
399
- 3. Generate plotHeatmap with appropriate visualization parameters
400
- 4. Suggest clustering if dataset is large
401
- 5. Offer profile plot as complement
402
-
403
- ## Key Reminders
404
-
405
- - **File validation first**: Always validate input files before analysis
406
- - **Normalization matters**: Choose appropriate method for comparison type
407
- - **Extend reads carefully**: YES for ChIP-seq, NO for RNA-seq
408
- - **Use all cores**: Set `--numberOfProcessors` to available cores
409
- - **Test on regions**: Use `--region` for parameter testing
410
- - **Check QC first**: Run quality control before detailed analysis
411
- - **Document everything**: Save commands for reproducibility
412
- - **Reference documentation**: Use comprehensive references for detailed guidance