@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,179 +0,0 @@
1
- ---
2
- name: pyopenms
3
- description: Complete mass spectrometry analysis platform. Use for proteomics and metabolomics workflows—feature detection, peptide/protein identification, label-free and isobaric quantification, adduct/accurate-mass annotation, and complex LC-MS/MS pipelines. Supports extensive file formats and algorithms. For simple spectral comparison and small-molecule library matching use matchms.
4
- license: 3 clause BSD license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.9+ and uv. Examples and scripts target pyOpenMS 3.5.0.
7
- metadata:
8
- version: "2.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # PyOpenMS
13
-
14
- ## Overview
15
-
16
- PyOpenMS provides Python bindings to the OpenMS library for computational mass
17
- spectrometry, enabling analysis of proteomics and metabolomics data. Use it to
18
- read/write MS file formats, process raw spectra, detect and quantify features,
19
- identify peptides and proteins, and run end-to-end LC-MS/MS pipelines.
20
-
21
- **This skill ships ready-to-run scripts in `scripts/`** covering the most common
22
- high-level workflows. Prefer running a script over writing new code—each is a
23
- parameterized CLI tool that handles loading, processing, and export. Drop into the
24
- Python API (and the `references/`) only when no script fits.
25
-
26
- ## Installation
27
-
28
- ```bash
29
- uv pip install pyopenms
30
- ```
31
-
32
- Verify (note: `__version__` works, but the bundled binary prints a one-line
33
- memory-status notice on import that is harmless):
34
-
35
- ```python
36
- import pyopenms as ms
37
- print(ms.__version__) # 3.5.0
38
- ```
39
-
40
- ## Scripts (start here)
41
-
42
- Run with `python scripts/<name>.py --help` for full options. All accept standard
43
- MS file formats and write featureXML/consensusXML/CSV/mzTab/PNG as appropriate.
44
-
45
- ### Inspect & convert
46
- | Script | What it does |
47
- |--------|--------------|
48
- | `inspect_ms_data.py` | Summarize any mzML/mzXML/featureXML/consensusXML/idXML (counts, RT/m/z ranges, TIC, metadata); optional per-spectrum CSV. |
49
- | `convert_format.py` | Convert between mzML/mzXML/MGF with optional MS-level, RT, and intensity filtering. |
50
- | `process_spectra.py` | Configurable signal-processing chain: smoothing (Gauss/SGolay), centroiding (PeakPickerHiRes), normalization, S/N and intensity thresholds. |
51
-
52
- ### Feature detection & quantification
53
- | Script | What it does |
54
- |--------|--------------|
55
- | `detect_features_metabo.py` | Untargeted metabolomics feature finding: MassTraceDetection → ElutionPeakDetection → FeatureFindingMetabo. |
56
- | `detect_features_centroided.py` | Peptide/centroided feature detection via FeatureFinderAlgorithmPicked. |
57
- | `align_link_quantify.py` | Multi-sample pipeline: detect (or load) features → RT alignment → consensus linking → quant matrix CSV. |
58
- | `consensus_to_matrix.py` | consensusXML → wide intensity matrix + metadata, with optional median/quantile normalization and long format. |
59
-
60
- ### Annotation
61
- | Script | What it does |
62
- |--------|--------------|
63
- | `detect_adducts.py` | Group adducts/charge variants of the same neutral mass (MetaboliteFeatureDeconvolution). |
64
- | `accurate_mass_search.py` | Annotate features against HMDB by accurate mass (AccurateMassSearchEngine → mzTab/CSV). |
65
- | `export_gnps_sirius.py` | Export GNPS FBMN inputs (MGF + quant table) or a SIRIUS `.ms` file. |
66
-
67
- ### Identification
68
- | Script | What it does |
69
- |--------|--------------|
70
- | `process_identifications.py` | Re-index against FASTA, estimate FDR/q-values, filter (FDR/length/best-per-spectrum), export idXML + CSV. |
71
-
72
- ### Chemistry
73
- | Script | What it does |
74
- |--------|--------------|
75
- | `mass_calculator.py` | Monoisotopic/average mass, charged m/z, formula, and isotope pattern for peptides or empirical formulas. |
76
- | `digest_protein.py` | In-silico protease digestion of FASTA/sequence → theoretical peptides with masses and m/z. |
77
- | `theoretical_spectrum.py` | Generate annotated theoretical fragment spectra (b/y/a/c/x/z, losses) for a peptide. |
78
-
79
- ### Targeted & visualization
80
- | Script | What it does |
81
- |--------|--------------|
82
- | `extract_chromatograms.py` | Build TIC/BPC and XIC traces for target m/z (CSV + optional plot). |
83
- | `plot_ms_data.py` | Quick plots: single spectrum, TIC, 2D feature map, MS1 signal map. |
84
-
85
- ### Common script recipes
86
-
87
- ```bash
88
- # Inspect a file
89
- python scripts/inspect_ms_data.py sample.mzML --spectra-csv spectra.csv
90
-
91
- # Untargeted metabolomics: features for one sample
92
- python scripts/detect_features_metabo.py sample.mzML --out-csv features.csv
93
-
94
- # Full multi-sample quantification study
95
- python scripts/align_link_quantify.py s1.mzML s2.mzML s3.mzML --out-prefix study
96
- python scripts/consensus_to_matrix.py study.consensusXML --out quant.csv --normalize median
97
-
98
- # Peptide chemistry
99
- python scripts/mass_calculator.py --peptide "PEPTIDEM(Oxidation)K" --charges 1 2 3 --isotopes 5
100
- python scripts/digest_protein.py proteins.fasta --enzyme Trypsin --missed 2 --out peptides.csv
101
-
102
- # Identification post-processing
103
- python scripts/process_identifications.py search.idXML --fasta db.fasta --fdr 0.01 --out filtered.idXML --csv hits.csv
104
- ```
105
-
106
- ## Key 3.5.0 API notes
107
-
108
- These changed from older OpenMS releases—older tutorials and code will break:
109
-
110
- - **Feature finding**: `FeatureFinder("centroided")` was **removed**. Use
111
- `FeatureFinderAlgorithmPicked` (proteomics/centroided) or the
112
- `MassTraceDetection → ElutionPeakDetection → FeatureFindingMetabo` pipeline
113
- (metabolomics). See `detect_features_*.py`.
114
- - **idXML I/O**: `IdXMLFile().load/store` require a `ms.PeptideIdentificationList()`
115
- for peptide IDs (a plain Python `list` raises "can not handle type"). Protein IDs
116
- remain a plain list.
117
- - **Adduct decharging**: the class is `MetaboliteFeatureDeconvolution`, and adducts
118
- use `Elements:Charge:Probability` syntax (e.g. `H:+:0.4`, `H-2O-1:0:0.05`)—not
119
- bracket notation like `[M+H]+`.
120
- - **DataFrame columns**: `FeatureMap.get_df()` uses lowercase `rt`/`mz` (not `RT`).
121
- `ConsensusMap` provides `get_intensity_df()` and `get_metadata_df()`.
122
- - **Bundled data caveat**: the pip wheel ships `HMDBMappingFile.tsv` but not
123
- `HMDB2StructMapping.tsv`; `accurate_mass_search.py` detects this and explains how
124
- to supply it.
125
-
126
- ## Core data structures
127
-
128
- - **MSExperiment** – collection of spectra and chromatograms
129
- - **MSSpectrum / MSChromatogram** – a single spectrum / chromatographic trace
130
- - **Feature / FeatureMap** – a detected LC-MS peak / collection of features
131
- - **ConsensusMap** – features linked across samples (the quant table)
132
- - **PeptideIdentification / ProteinIdentification** – search results
133
- - **AASequence / EmpiricalFormula** – sequence and formula chemistry
134
-
135
- **For details**: see `references/data_structures.md`.
136
-
137
- ## Parameter management
138
-
139
- Most algorithms expose an OpenMS `Param` object:
140
-
141
- ```python
142
- algo = ms.FeatureFindingMetabo()
143
- p = algo.getDefaults()
144
- for key in p.keys():
145
- print(key.decode(), "=", p.getValue(key), "|", p.getDescription(key))
146
- p.setValue("charge_lower_bound", 1)
147
- algo.setParameters(p)
148
- ```
149
-
150
- ## Export to pandas
151
-
152
- ```python
153
- fm = ms.FeatureMap(); ms.FeatureXMLFile().load("features.featureXML", fm)
154
- df = fm.get_df() # columns include lowercase rt, mz, intensity, charge, quality
155
-
156
- cm = ms.ConsensusMap(); ms.ConsensusXMLFile().load("study.consensusXML", cm)
157
- intensities = cm.get_intensity_df() # features x samples
158
- metadata = cm.get_metadata_df() # rt, mz, charge, quality, ...
159
- ```
160
-
161
- ## Integration with other tools
162
-
163
- Pandas (DataFrames), NumPy (peak arrays), scikit-learn (ML), Matplotlib/Seaborn
164
- (plots), and downstream tools via export: GNPS (FBMN), SIRIUS, and mzTab.
165
-
166
- ## Resources
167
-
168
- - Official docs (3.5.0): https://pyopenms.readthedocs.io/en/release-3.5.0/
169
- - OpenMS: https://www.openms.org
170
- - GitHub: https://github.com/OpenMS/OpenMS
171
-
172
- ## References
173
-
174
- - `references/file_io.md` – file format handling
175
- - `references/signal_processing.md` – signal processing algorithms
176
- - `references/feature_detection.md` – feature detection and linking
177
- - `references/identification.md` – peptide and protein identification
178
- - `references/metabolomics.md` – metabolomics-specific workflows
179
- - `references/data_structures.md` – core objects and data structures
@@ -1,330 +0,0 @@
1
- ---
2
- name: pysam
3
- description: Python/HTSlib workflows for genomic files. Use when reading, querying, filtering, or writing SAM/BAM/CRAM, VCF/BCF, FASTA/FASTQ, or tabix data with pysam, including pileup, coverage, indexing, and CRAM references.
4
- license: MIT
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.8–3.14 and pysam 0.24.0. Bundled scripts use local files. CRAM decoding may require the matching reference FASTA or an explicitly configured REF_PATH/REF_CACHE.
7
- metadata:
8
- version: "2.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # pysam
13
-
14
- ## Overview
15
-
16
- Use pysam for low-level, streaming access to HTSlib-supported genomic formats:
17
-
18
- - `AlignmentFile` and `AlignedSegment` for SAM/BAM/CRAM
19
- - `VariantFile`, `VariantHeader`, and `VariantRecord` for VCF/BCF
20
- - `FastaFile` for indexed FASTA and `FastxFile` for sequential FASTA/FASTQ
21
- - `TabixFile` for BGZF-compressed, tabix-indexed BED/GFF/GTF/custom tables
22
- - `pysam.samtools` and `pysam.bcftools` for wrapped command dispatchers
23
-
24
- Current upstream baseline: **pysam 0.24.0** (27 April 2026), wrapping
25
- HTSlib/samtools/bcftools 1.23.1. Read `references/sources.md` before updating
26
- version-specific guidance.
27
-
28
- ## Installation
29
-
30
- Use the pinned release for reproducible work:
31
-
32
- ```bash
33
- uv pip install "pysam==0.24.0"
34
- ```
35
-
36
- Confirm the runtime:
37
-
38
- ```python
39
- import pysam
40
-
41
- print(pysam.__version__) # 0.24.0
42
- print(pysam.__samtools_version__) # 1.23.1
43
- ```
44
-
45
- Prebuilt wheels are available for supported macOS and Linux platforms. A
46
- source build needs a C compiler and HTSlib build dependencies; read the
47
- official installation guide linked from `references/sources.md`.
48
-
49
- ## First Decide
50
-
51
- Before writing code:
52
-
53
- 1. Identify the real format, compression, sort order, and available index.
54
- 2. Decide whether coordinates are numeric Python coordinates or a region
55
- string. Do not mix them.
56
- 3. For CRAM, identify the exact reference assembly and FASTA.
57
- 4. Prefer indexed region access; use sequential iteration only when intended.
58
- 5. Preserve headers when writing and write to a new path by default.
59
- 6. State filtering semantics: mapping/base quality, flags, overlap handling,
60
- duplicate handling, and pileup depth cap.
61
-
62
- For unfamiliar files, start with the bundled read-only inspector:
63
-
64
- ```bash
65
- python scripts/inspect_hts.py sample.bam
66
- python scripts/inspect_hts.py cohort.vcf.gz
67
- python scripts/inspect_hts.py reference.fa
68
- ```
69
-
70
- ## Bundled Scripts
71
-
72
- | Script | Purpose | Typical call |
73
- |---|---|---|
74
- | `scripts/inspect_hts.py` | Metadata-only inspection for alignment, variant, FASTA, FASTQ, and tabix files | `python scripts/inspect_hts.py sample.cram --reference ref.fa` |
75
- | `scripts/alignment_qc.py` | Streaming aggregate read/QC counts as JSON | `python scripts/alignment_qc.py sample.bam --max-records 100000` |
76
- | `scripts/variant_summary.py` | Streaming variant, FILTER, and genotype summary as JSON | `python scripts/variant_summary.py cohort.vcf.gz --region chr1:1-1000000` |
77
- | `scripts/filter_alignments.py` | Filter SAM/BAM/CRAM without changing record order | `python scripts/filter_alignments.py input.bam output.bam --exclude-secondary` |
78
-
79
- All scripts refuse to overwrite existing outputs. Run each with `--help` for
80
- coordinate, index, and privacy notes.
81
-
82
- ## Coordinate Contract
83
-
84
- **Numeric coordinates accepted by pysam APIs are 0-based, half-open.** This
85
- includes numeric `AlignmentFile.fetch()`, `VariantFile.fetch()`,
86
- `FastaFile.fetch()`, `TabixFile.fetch()`, and `pileup()` arguments.
87
-
88
- **Region strings are samtools-style: 1-based and inclusive.**
89
-
90
- ```python
91
- # The same 100 bases:
92
- bam.fetch("chr1", 99, 199) # [99, 199)
93
- bam.fetch(region="chr1:100-199") # 1-based inclusive
94
- ```
95
-
96
- VCF text uses 1-based `POS`, while record properties expose both systems:
97
-
98
- ```python
99
- record.pos # 1-based
100
- record.start # 0-based inclusive
101
- record.stop # 0-based exclusive
102
- ```
103
-
104
- Read `references/coordinates_and_indexing.md` for format conversions, overlap
105
- semantics, index choices, and contig-name checks.
106
-
107
- ## Alignment Files
108
-
109
- Use context managers and explicit modes:
110
-
111
- ```python
112
- import pysam
113
-
114
- with pysam.AlignmentFile("sample.bam", "rb", threads=4) as bam:
115
- for read in bam.fetch("chr1", 1_000, 2_000):
116
- if (
117
- not read.is_unmapped
118
- and not read.is_secondary
119
- and not read.is_supplementary
120
- and read.mapping_quality >= 30
121
- ):
122
- print(read.query_name, read.reference_start, read.cigarstring)
123
- ```
124
-
125
- Use `fetch(until_eof=True)` to stream every record in file order, including
126
- unplaced unmapped reads, without requiring an index:
127
-
128
- ```python
129
- with pysam.AlignmentFile("sample.bam", "rb") as bam:
130
- for read in bam.fetch(until_eof=True):
131
- ...
132
- ```
133
-
134
- Important distinctions:
135
-
136
- - `fetch()` returns alignment records overlapping a region.
137
- - `count()` counts records and defaults to `read_callback="nofilter"`.
138
- - `count_coverage()` returns A/C/G/T base counts and defaults to base quality
139
- 15 plus `read_callback="all"`.
140
- - `pileup()` exposes per-column reads and has its own filtering, base-quality,
141
- overlap, orphan, and `max_depth=8000` defaults.
142
-
143
- For exact-region pileups, set `truncate=True` and explicit filters:
144
-
145
- ```python
146
- with pysam.FastaFile("reference.fa") as fasta, pysam.AlignmentFile(
147
- "sample.bam", "rb"
148
- ) as bam:
149
- for column in bam.pileup(
150
- "chr1",
151
- 1_000,
152
- 2_000,
153
- truncate=True,
154
- stepper="samtools",
155
- fastafile=fasta,
156
- min_mapping_quality=20,
157
- min_base_quality=20,
158
- max_depth=100_000,
159
- ):
160
- print(column.reference_pos, column.get_num_aligned())
161
- ```
162
-
163
- Read `references/alignment_files.md` for flags, CIGAR operations, tags,
164
- modified bases, writing records, pileup details, and iterator lifetime.
165
-
166
- ## Variant Files
167
-
168
- Input format is auto-detected. Numeric fetch coordinates remain 0-based:
169
-
170
- ```python
171
- import pysam
172
-
173
- with pysam.VariantFile("cohort.vcf.gz", threads=4) as variants:
174
- for record in variants.fetch("chr1", 999_999, 2_000_000):
175
- print(record.contig, record.pos, record.ref, record.alts)
176
- for sample_name, call in record.samples.items():
177
- print(sample_name, call.get("GT"))
178
- ```
179
-
180
- Subset samples **before retrieving records**:
181
-
182
- ```python
183
- with pysam.VariantFile("cohort.bcf") as variants:
184
- variants.subset_samples(["sample_A", "sample_B"])
185
- for record in variants:
186
- ...
187
- ```
188
-
189
- When changing a header, copy each record and translate it to the destination
190
- header before assigning newly declared INFO/FORMAT/FILTER fields. Do not
191
- manually clear and rebuild `header.samples`.
192
-
193
- Read `references/variant_files.md` for safe headers, writing, sample
194
- subsetting, missing genotypes, symbolic alleles, filtering, translation, and
195
- indexing.
196
-
197
- ## FASTA, FASTQ, and Tabix
198
-
199
- Indexed FASTA uses numeric 0-based coordinates:
200
-
201
- ```python
202
- with pysam.FastaFile("reference.fa") as fasta:
203
- sequence = fasta.fetch("chr1", 999, 1_099)
204
- ```
205
-
206
- `FastxFile` is sequential. `persist=False` is faster but yielded records become
207
- invalid after iteration advances:
208
-
209
- ```python
210
- with pysam.FastxFile("reads.fastq.gz", persist=False) as reads:
211
- for read in reads:
212
- qualities = read.get_quality_array()
213
- ...
214
- ```
215
-
216
- Tabix input must be coordinate-sorted and BGZF-compressed, not ordinary gzip.
217
- Use a non-destructive two-step workflow:
218
-
219
- ```python
220
- pysam.tabix_compress("regions.bed", "regions.bed.gz")
221
- pysam.tabix_index("regions.bed.gz", preset="bed")
222
-
223
- with pysam.TabixFile("regions.bed.gz", parser=pysam.asBed()) as tbx:
224
- for interval in tbx.fetch("chr1", 1_000, 2_000):
225
- print(interval.contig, interval.start, interval.end)
226
- ```
227
-
228
- Read `references/sequence_files.md` for FASTA/FASTQ records and safe tabix
229
- creation.
230
-
231
- ## CRAM, Remote I/O, and Threads
232
-
233
- pysam 0.24 changed inherited HTSlib behavior:
234
-
235
- - Newly written CRAM defaults to CRAM 3.1, not 3.0.
236
- - HTSlib no longer contacts the EBI reference server by default.
237
- - Prefer `reference_filename="reference.fa"` for deterministic local reads and
238
- writes.
239
-
240
- ```python
241
- with pysam.AlignmentFile(
242
- "sample.cram",
243
- "rc",
244
- reference_filename="reference.fa",
245
- threads=4,
246
- ) as cram:
247
- for read in cram.fetch("chr1", 1_000, 2_000):
248
- ...
249
- ```
250
-
251
- Only configure `REF_PATH`/`REF_CACHE` when reference-by-MD5 lookup is
252
- intentional. Do not assume a CRAM is self-contained. `threads=` accelerates
253
- compression/decompression; it does not parallelize Python analysis.
254
-
255
- Read `references/cram_and_performance.md` before CRAM conversion, remote access,
256
- or concurrent iteration.
257
-
258
- ## Wrapped samtools and bcftools
259
-
260
- Import command modules explicitly. Pass each command-line token as a separate
261
- string:
262
-
263
- ```python
264
- import pysam.samtools
265
- import pysam.bcftools
266
-
267
- pysam.samtools.sort(
268
- "-@", "4", "-o", "sorted.bam", "input.bam", catch_stdout=False
269
- )
270
- pysam.samtools.index("-@", "4", "sorted.bam", catch_stdout=False)
271
-
272
- pysam.bcftools.index("--csi", "variants.vcf.gz", catch_stdout=False)
273
- ```
274
-
275
- Dispatchers capture stdout by default. For large or binary output, use the
276
- tool's `-o` option with `catch_stdout=False`, or `save_stdout=...`, rather than
277
- returning the complete output in memory.
278
-
279
- ```python
280
- try:
281
- pysam.samtools.quickcheck("-v", "sample.bam")
282
- except pysam.SamtoolsError as error:
283
- messages = pysam.samtools.quickcheck.get_messages()
284
- raise RuntimeError(messages or str(error)) from error
285
- ```
286
-
287
- Use the Python API for record-level logic and dispatchers for mature bulk
288
- operations such as sort, index, merge, view, and normalization. Never compose
289
- dispatcher arguments by splitting an untrusted shell command.
290
-
291
- ## Writing Rules
292
-
293
- - Copy or construct a valid header before opening output.
294
- - Write to a new path; do not use `force=True` unless replacement is explicit.
295
- - Preserve sort order if the output will be indexed.
296
- - Set `query_sequence` before `query_qualities`.
297
- - Prefer `pysam.CIGAR_OPS` enum members; top-level constants such as
298
- `pysam.CMATCH` are compatibility aliases slated for future removal.
299
- - Validate outputs with `pysam.samtools.quickcheck()` for alignments and reopen
300
- variant/sequence outputs before downstream use.
301
- - Use CSI rather than BAI/TBI when references or coordinates exceed legacy
302
- index limits.
303
-
304
- ## Reference Map
305
-
306
- | Need | Read |
307
- |---|---|
308
- | Alignment API, flags, CIGAR, pileup, modified bases | `references/alignment_files.md` |
309
- | VCF/BCF headers, records, samples, writing | `references/variant_files.md` |
310
- | FASTA/FASTQ and tabix-indexed tables | `references/sequence_files.md` |
311
- | Coordinate conversion and index selection | `references/coordinates_and_indexing.md` |
312
- | CRAM references, remote I/O, threads, performance | `references/cram_and_performance.md` |
313
- | Correct integrated analysis patterns | `references/common_workflows.md` |
314
- | Compact current API signatures and defaults | `references/api_reference.md` |
315
- | Upgrade notes for existing environments | `references/migration_to_0_24.md` |
316
- | Official docs, specifications, and release sources | `references/sources.md` |
317
-
318
- ## Common Failure Modes
319
-
320
- - Treating numeric `VariantFile.fetch()` coordinates as 1-based
321
- - Using ordinary gzip where BGZF plus tabix/CSI is required
322
- - Calling region fetch without an index
323
- - Assuming `fetch()` includes unplaced unmapped alignments
324
- - Forgetting `truncate=True` for an exact pileup interval
325
- - Ignoring pileup defaults such as base quality 13 and depth cap 8000
326
- - Sharing one file handle across active iterators or threads
327
- - Decoding CRAM without its exact reference
328
- - Assigning a new VCF field before declaring it in the output header
329
- - Capturing large samtools/bcftools output in memory
330
- - Using a SNP base-counting method for indels or symbolic alleles