@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,379 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: polars-bio
|
|
3
|
-
description: High-performance genomic interval operations and bioinformatics file I/O on Polars DataFrames. Overlap, nearest, merge, coverage, complement, subtract for BED/VCF/BAM/GFF intervals. Streaming, cloud-native, faster bioframe alternative.
|
|
4
|
-
license: Apache-2.0
|
|
5
|
-
allowed-tools: Read Write Edit Bash
|
|
6
|
-
compatibility: Requires Python 3.11–3.14 and polars-bio (uv pip install). Cloud I/O uses standard AWS/GCS/Azure SDK env vars when paths use s3://, gs://, or az:// URIs.
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.0"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# polars-bio
|
|
13
|
-
|
|
14
|
-
## Overview
|
|
15
|
-
|
|
16
|
-
polars-bio is a high-performance Python library for genomic interval operations and bioinformatics file I/O, built on Polars, Apache Arrow, and Apache DataFusion. It provides a familiar DataFrame-centric API for interval arithmetic (overlap, nearest, merge, coverage, complement, subtract) and reading/writing common bioinformatics formats (BED, VCF, BAM, CRAM, GFF/GTF, FASTA, FASTQ).
|
|
17
|
-
|
|
18
|
-
Key value propositions:
|
|
19
|
-
- **6-38x faster** than bioframe on real-world genomic benchmarks
|
|
20
|
-
- **Streaming/out-of-core** support for large genomes via DataFusion
|
|
21
|
-
- **Cloud-native** file I/O (S3, GCS, Azure) with predicate pushdown
|
|
22
|
-
- **Two API styles**: functional (`pb.overlap(df1, df2)`) and method-chaining (`df1.lazy().pb.overlap(df2)`)
|
|
23
|
-
- **SQL interface** for genomic data via DataFusion SQL engine
|
|
24
|
-
|
|
25
|
-
## When to Use This Skill
|
|
26
|
-
|
|
27
|
-
Use this skill when:
|
|
28
|
-
- Performing genomic interval operations (overlap, nearest, merge, coverage, complement, subtract)
|
|
29
|
-
- Reading/writing bioinformatics file formats (BED, VCF, BAM, CRAM, GFF/GTF, FASTA, FASTQ)
|
|
30
|
-
- Processing large genomic datasets that don't fit in memory (streaming mode)
|
|
31
|
-
- Running SQL queries on genomic data files
|
|
32
|
-
- Migrating from bioframe to a faster alternative
|
|
33
|
-
- Computing read depth/pileup from BAM/CRAM files
|
|
34
|
-
- Working with Polars DataFrames containing genomic intervals
|
|
35
|
-
|
|
36
|
-
## Quick Start
|
|
37
|
-
|
|
38
|
-
### Installation
|
|
39
|
-
|
|
40
|
-
Requires Python 3.11–3.14 (see [PyPI](https://pypi.org/project/polars-bio/)).
|
|
41
|
-
|
|
42
|
-
```bash
|
|
43
|
-
uv pip install "polars-bio==0.31.0"
|
|
44
|
-
```
|
|
45
|
-
|
|
46
|
-
For pandas compatibility (pandas ≥3.0):
|
|
47
|
-
|
|
48
|
-
```bash
|
|
49
|
-
uv pip install "polars-bio[pandas]==0.31.0"
|
|
50
|
-
```
|
|
51
|
-
|
|
52
|
-
### Basic Overlap Example
|
|
53
|
-
|
|
54
|
-
```python
|
|
55
|
-
import polars as pl
|
|
56
|
-
import polars_bio as pb
|
|
57
|
-
|
|
58
|
-
# Create two interval DataFrames
|
|
59
|
-
df1 = pl.DataFrame({
|
|
60
|
-
"chrom": ["chr1", "chr1", "chr1"],
|
|
61
|
-
"start": [1, 5, 22],
|
|
62
|
-
"end": [6, 9, 30],
|
|
63
|
-
})
|
|
64
|
-
|
|
65
|
-
df2 = pl.DataFrame({
|
|
66
|
-
"chrom": ["chr1", "chr1"],
|
|
67
|
-
"start": [3, 25],
|
|
68
|
-
"end": [8, 28],
|
|
69
|
-
})
|
|
70
|
-
|
|
71
|
-
# Functional API (returns LazyFrame by default)
|
|
72
|
-
result = pb.overlap(df1, df2)
|
|
73
|
-
result_df = result.collect()
|
|
74
|
-
|
|
75
|
-
# Get a DataFrame directly
|
|
76
|
-
result_df = pb.overlap(df1, df2, output_type="polars.DataFrame")
|
|
77
|
-
|
|
78
|
-
# Method-chaining API (via .pb accessor on LazyFrame)
|
|
79
|
-
result = df1.lazy().pb.overlap(df2)
|
|
80
|
-
result_df = result.collect()
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
### Reading a BED File
|
|
84
|
-
|
|
85
|
-
```python
|
|
86
|
-
import polars_bio as pb
|
|
87
|
-
|
|
88
|
-
# Eager read (loads entire file)
|
|
89
|
-
df = pb.read_bed("regions.bed")
|
|
90
|
-
|
|
91
|
-
# Lazy scan (streaming, for large files)
|
|
92
|
-
lf = pb.scan_bed("regions.bed")
|
|
93
|
-
result = lf.collect()
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
## Core Capabilities
|
|
97
|
-
|
|
98
|
-
### 1. Genomic Interval Operations
|
|
99
|
-
|
|
100
|
-
polars-bio provides 8 core interval operations for genomic range arithmetic. All operations accept Polars DataFrames with `chrom`, `start`, `end` columns (configurable). All operations return a `LazyFrame` by default (use `output_type="polars.DataFrame"` for eager results).
|
|
101
|
-
|
|
102
|
-
**Operations:**
|
|
103
|
-
- `overlap` / `count_overlaps` - Find or count overlapping intervals between two sets (`overlap_output="left"` returns df1-only hits since 0.30.0)
|
|
104
|
-
- `nearest` - Find nearest intervals (with configurable `k`, `overlap`, `distance` params)
|
|
105
|
-
- `merge` - Merge overlapping/bookended intervals within a set
|
|
106
|
-
- `cluster` - Assign cluster IDs to overlapping intervals
|
|
107
|
-
- `coverage` - Compute per-interval coverage counts (two-input operation)
|
|
108
|
-
- `complement` - Find gaps between intervals within a genome
|
|
109
|
-
- `subtract` - Remove portions of intervals that overlap another set
|
|
110
|
-
|
|
111
|
-
**Example:**
|
|
112
|
-
```python
|
|
113
|
-
import polars_bio as pb
|
|
114
|
-
|
|
115
|
-
# Find overlapping intervals (returns LazyFrame)
|
|
116
|
-
result = pb.overlap(df1, df2, suffixes=("_1", "_2"))
|
|
117
|
-
|
|
118
|
-
# Count overlaps per interval
|
|
119
|
-
counts = pb.count_overlaps(df1, df2)
|
|
120
|
-
|
|
121
|
-
# Merge overlapping intervals
|
|
122
|
-
merged = pb.merge(df1)
|
|
123
|
-
|
|
124
|
-
# Find nearest intervals
|
|
125
|
-
nearest = pb.nearest(df1, df2)
|
|
126
|
-
|
|
127
|
-
# Collect any LazyFrame result to DataFrame
|
|
128
|
-
result_df = result.collect()
|
|
129
|
-
```
|
|
130
|
-
|
|
131
|
-
**Reference:** See `references/interval_operations.md` for detailed documentation on all operations, parameters, output schemas, and performance considerations.
|
|
132
|
-
|
|
133
|
-
### 2. Bioinformatics File I/O
|
|
134
|
-
|
|
135
|
-
Read and write common bioinformatics formats with `read_*`, `scan_*`, `write_*`, and `sink_*` functions. Supports cloud storage (S3, GCS, Azure) and compression (GZIP, BGZF).
|
|
136
|
-
|
|
137
|
-
**Supported formats:**
|
|
138
|
-
- **BED** - Genomic intervals (`read_bed`, `scan_bed`, `write_*` via generic)
|
|
139
|
-
- **VCF** - Genetic variants (`read_vcf`, `scan_vcf`, `write_vcf`, `sink_vcf`)
|
|
140
|
-
- **VCF Zarr** - Analysis-ready Zarr stores (`read_vcf_zarr`, `scan_vcf_zarr`; local directory paths)
|
|
141
|
-
- **BAM** - Aligned reads (`read_bam`, `scan_bam`, `write_bam`, `sink_bam`)
|
|
142
|
-
- **CRAM** - Compressed alignments (`read_cram`, `scan_cram`, `write_cram`, `sink_cram`)
|
|
143
|
-
- **GFF** - Gene annotations (`read_gff`, `scan_gff`)
|
|
144
|
-
- **GTF** - Gene annotations (`read_gtf`, `scan_gtf`)
|
|
145
|
-
- **FASTA** - Reference sequences (`read_fasta`, `scan_fasta`, `write_fasta`, `sink_fasta`)
|
|
146
|
-
- **FASTQ** - Sequencing reads (`read_fastq`, `scan_fastq`, `write_fastq`, `sink_fastq`)
|
|
147
|
-
- **SAM** - Text alignments (`read_sam`, `scan_sam`, `write_sam`, `sink_sam`)
|
|
148
|
-
- **Hi-C pairs** - Chromatin contacts (`read_pairs`, `scan_pairs`)
|
|
149
|
-
|
|
150
|
-
**Example:**
|
|
151
|
-
```python
|
|
152
|
-
import polars_bio as pb
|
|
153
|
-
|
|
154
|
-
# Read VCF file
|
|
155
|
-
variants = pb.read_vcf("samples.vcf.gz")
|
|
156
|
-
|
|
157
|
-
# Lazy scan BAM file (streaming)
|
|
158
|
-
alignments = pb.scan_bam("aligned.bam")
|
|
159
|
-
|
|
160
|
-
# Read GFF annotations
|
|
161
|
-
genes = pb.read_gff("annotations.gff3")
|
|
162
|
-
|
|
163
|
-
# Cloud storage (individual params, not a dict)
|
|
164
|
-
df = pb.read_bed("s3://bucket/regions.bed",
|
|
165
|
-
allow_anonymous=True)
|
|
166
|
-
```
|
|
167
|
-
|
|
168
|
-
**Reference:** See `references/file_io.md` for per-format column schemas, parameters, cloud storage options, and compression support.
|
|
169
|
-
|
|
170
|
-
### 3. SQL Data Processing
|
|
171
|
-
|
|
172
|
-
Register bioinformatics files as tables and query them using DataFusion SQL. Combines the power of SQL with polars-bio's genomic-aware readers.
|
|
173
|
-
|
|
174
|
-
```python
|
|
175
|
-
import polars as pl
|
|
176
|
-
import polars_bio as pb
|
|
177
|
-
|
|
178
|
-
# Register files as SQL tables (path first, name= keyword)
|
|
179
|
-
pb.register_vcf("samples.vcf.gz", name="variants")
|
|
180
|
-
pb.register_bed("target_regions.bed", name="regions")
|
|
181
|
-
|
|
182
|
-
# Query with SQL (returns LazyFrame)
|
|
183
|
-
result = pb.sql("SELECT chrom, start, end, ref, alt FROM variants WHERE qual > 30")
|
|
184
|
-
result_df = result.collect()
|
|
185
|
-
|
|
186
|
-
# Register a Polars DataFrame as a SQL table
|
|
187
|
-
pb.from_polars("my_intervals", df)
|
|
188
|
-
result = pb.sql("SELECT * FROM my_intervals WHERE chrom = 'chr1'").collect()
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
**Reference:** See `references/sql_processing.md` for register functions, SQL syntax, and examples.
|
|
192
|
-
|
|
193
|
-
### 4. Pileup Operations
|
|
194
|
-
|
|
195
|
-
Compute per-base read depth from BAM/CRAM files with CIGAR-aware depth calculation.
|
|
196
|
-
|
|
197
|
-
```python
|
|
198
|
-
import polars_bio as pb
|
|
199
|
-
|
|
200
|
-
# Compute depth across a BAM file
|
|
201
|
-
depth_lf = pb.depth("aligned.bam")
|
|
202
|
-
depth_df = depth_lf.collect()
|
|
203
|
-
|
|
204
|
-
# With quality filter
|
|
205
|
-
depth_lf = pb.depth("aligned.bam", min_mapping_quality=20)
|
|
206
|
-
```
|
|
207
|
-
|
|
208
|
-
**Reference:** See `references/pileup_operations.md` for parameters and integration patterns.
|
|
209
|
-
|
|
210
|
-
## Key Concepts
|
|
211
|
-
|
|
212
|
-
### Coordinate Systems
|
|
213
|
-
|
|
214
|
-
polars-bio defaults to **1-based** coordinates (genomic convention). This can be changed globally:
|
|
215
|
-
|
|
216
|
-
```python
|
|
217
|
-
import polars_bio as pb
|
|
218
|
-
|
|
219
|
-
# Switch to 0-based half-open coordinates (default is 1-based / False)
|
|
220
|
-
pb.set_option("datafusion.bio.coordinate_system_zero_based", True)
|
|
221
|
-
|
|
222
|
-
# Switch back to 1-based (default)
|
|
223
|
-
pb.set_option("datafusion.bio.coordinate_system_zero_based", False)
|
|
224
|
-
```
|
|
225
|
-
|
|
226
|
-
I/O functions also accept `use_zero_based` to set coordinate metadata on the resulting DataFrame:
|
|
227
|
-
|
|
228
|
-
```python
|
|
229
|
-
# Read BED with explicit 0-based metadata
|
|
230
|
-
df = pb.read_bed("regions.bed", use_zero_based=True)
|
|
231
|
-
```
|
|
232
|
-
|
|
233
|
-
**Important:** BED files are always 0-based half-open in the file format. polars-bio handles the conversion automatically when reading BED files. Coordinate metadata is attached to DataFrames by I/O functions and propagated through operations.
|
|
234
|
-
|
|
235
|
-
### Two API Styles
|
|
236
|
-
|
|
237
|
-
**Functional API** - standalone functions, explicit inputs:
|
|
238
|
-
```python
|
|
239
|
-
result = pb.overlap(df1, df2, suffixes=("_1", "_2"))
|
|
240
|
-
merged = pb.merge(df)
|
|
241
|
-
```
|
|
242
|
-
|
|
243
|
-
**Method-chaining API** - via `.pb` accessor on **LazyFrames** (not DataFrames):
|
|
244
|
-
```python
|
|
245
|
-
result = df1.lazy().pb.overlap(df2)
|
|
246
|
-
merged = df.lazy().pb.merge()
|
|
247
|
-
```
|
|
248
|
-
|
|
249
|
-
**Important:** The `.pb` accessor for interval operations is only available on `LazyFrame`. On `DataFrame`, `.pb` provides write operations only (`write_bam`, `write_vcf`, etc.).
|
|
250
|
-
|
|
251
|
-
Method-chaining enables fluent pipelines:
|
|
252
|
-
```python
|
|
253
|
-
# Chain interval operations (note: overlap outputs suffixed columns,
|
|
254
|
-
# so rename before merge which expects chrom/start/end)
|
|
255
|
-
result = (
|
|
256
|
-
df1.lazy()
|
|
257
|
-
.pb.overlap(df2)
|
|
258
|
-
.filter(pl.col("start_2") > 1000)
|
|
259
|
-
.select(
|
|
260
|
-
pl.col("chrom_1").alias("chrom"),
|
|
261
|
-
pl.col("start_1").alias("start"),
|
|
262
|
-
pl.col("end_1").alias("end"),
|
|
263
|
-
)
|
|
264
|
-
.pb.merge()
|
|
265
|
-
.collect()
|
|
266
|
-
)
|
|
267
|
-
```
|
|
268
|
-
|
|
269
|
-
### Probe-Build Architecture
|
|
270
|
-
|
|
271
|
-
For two-input operations (overlap, nearest, count_overlaps, coverage), polars-bio uses a probe-build join strategy:
|
|
272
|
-
- The **first** DataFrame is the **probe** (iterated over)
|
|
273
|
-
- The **second** DataFrame is the **build** (indexed for lookup)
|
|
274
|
-
|
|
275
|
-
For best performance, pass the larger DataFrame as the first argument (probe) and the smaller one as the second (build).
|
|
276
|
-
|
|
277
|
-
### Column Conventions
|
|
278
|
-
|
|
279
|
-
By default, polars-bio expects columns named `chrom`, `start`, `end`. Custom column names can be specified via lists:
|
|
280
|
-
|
|
281
|
-
```python
|
|
282
|
-
result = pb.overlap(
|
|
283
|
-
df1, df2,
|
|
284
|
-
cols1=["chromosome", "begin", "finish"],
|
|
285
|
-
cols2=["chr", "pos_start", "pos_end"],
|
|
286
|
-
)
|
|
287
|
-
```
|
|
288
|
-
|
|
289
|
-
### Return Types and Collecting Results
|
|
290
|
-
|
|
291
|
-
All interval operations and `pb.sql()` return a **LazyFrame** by default. Use `.collect()` to materialize results, or pass `output_type="polars.DataFrame"` for eager evaluation:
|
|
292
|
-
|
|
293
|
-
```python
|
|
294
|
-
# Lazy (default) - collect when needed
|
|
295
|
-
result_lf = pb.overlap(df1, df2)
|
|
296
|
-
result_df = result_lf.collect()
|
|
297
|
-
|
|
298
|
-
# Eager - get DataFrame directly
|
|
299
|
-
result_df = pb.overlap(df1, df2, output_type="polars.DataFrame")
|
|
300
|
-
```
|
|
301
|
-
|
|
302
|
-
### Streaming and Out-of-Core Processing
|
|
303
|
-
|
|
304
|
-
For datasets larger than available RAM, use `scan_*` functions and streaming execution:
|
|
305
|
-
|
|
306
|
-
```python
|
|
307
|
-
# Scan files lazily
|
|
308
|
-
lf = pb.scan_bed("large_intervals.bed")
|
|
309
|
-
|
|
310
|
-
# Process with Polars streaming (requires polars ≥1.37, bundled with polars-bio)
|
|
311
|
-
result = lf.collect(engine="streaming")
|
|
312
|
-
```
|
|
313
|
-
|
|
314
|
-
DataFusion streaming is enabled by default for interval operations, processing data in batches without loading the full dataset into memory.
|
|
315
|
-
|
|
316
|
-
## Common Pitfalls
|
|
317
|
-
|
|
318
|
-
1. **`.pb` accessor on DataFrame vs LazyFrame:** Interval operations (overlap, merge, etc.) are only on `LazyFrame.pb`. `DataFrame.pb` only has write methods. Use `.lazy()` to convert before chaining interval ops.
|
|
319
|
-
|
|
320
|
-
2. **LazyFrame returns:** All interval operations and `pb.sql()` return `LazyFrame` by default. Don't forget `.collect()` or use `output_type="polars.DataFrame"`.
|
|
321
|
-
|
|
322
|
-
3. **Column name mismatches:** polars-bio expects `chrom`, `start`, `end` by default. Use `cols1`/`cols2` parameters (as lists) if your columns have different names.
|
|
323
|
-
|
|
324
|
-
4. **Coordinate system metadata:** Interval operations read coordinate metadata from I/O functions or DataFrame `config_meta`. For manually built DataFrames, set `df.config_meta.set(coordinate_system_zero_based=True)` (0-based) or `False` (1-based). If metadata is missing, polars-bio falls back to the global `datafusion.bio.coordinate_system_zero_based` setting (with a warning). Set `pb.set_option("datafusion.bio.coordinate_system_check", True)` to raise `MissingCoordinateSystemError` instead. Mismatched systems between inputs raise `CoordinateSystemMismatchError`.
|
|
325
|
-
|
|
326
|
-
5. **Probe-build order matters:** For overlap, nearest, and coverage, the first DataFrame is probed against the second. Swapping arguments changes which intervals appear in the left vs right output columns, and can affect performance.
|
|
327
|
-
|
|
328
|
-
6. **INT32 position limit:** Genomic positions are stored as 32-bit integers, limiting coordinates to ~2.1 billion. This is sufficient for all known genomes but may be an issue with custom coordinate spaces.
|
|
329
|
-
|
|
330
|
-
7. **BAM index requirements:** `read_bam` and `scan_bam` require a `.bai` index file alongside the BAM. Create one with `samtools index` if missing.
|
|
331
|
-
|
|
332
|
-
8. **Parallel execution disabled by default:** DataFusion parallelism defaults to 1 partition. Enable for large datasets:
|
|
333
|
-
```python
|
|
334
|
-
pb.set_option("datafusion.execution.target_partitions", 8)
|
|
335
|
-
```
|
|
336
|
-
|
|
337
|
-
9. **CRAM has separate functions:** Use `read_cram`/`scan_cram`/`register_cram` for CRAM files (not `read_bam`). CRAM functions require a `reference_path` parameter.
|
|
338
|
-
|
|
339
|
-
## Best Practices
|
|
340
|
-
|
|
341
|
-
1. **Use `scan_*` for large files:** Prefer `scan_bed`, `scan_vcf`, etc. over `read_*` for files larger than available RAM. Scan functions enable streaming and predicate pushdown.
|
|
342
|
-
|
|
343
|
-
2. **Configure parallelism for large datasets:**
|
|
344
|
-
```python
|
|
345
|
-
import os
|
|
346
|
-
pb.set_option("datafusion.execution.target_partitions", os.cpu_count())
|
|
347
|
-
```
|
|
348
|
-
|
|
349
|
-
3. **Use BGZF compression:** BGZF-compressed files (`.bed.gz`, `.vcf.gz`) support parallel block decompression, significantly faster than plain GZIP.
|
|
350
|
-
|
|
351
|
-
4. **Select columns early:** When only specific columns are needed, select them early to reduce memory usage:
|
|
352
|
-
```python
|
|
353
|
-
df = pb.read_vcf("large.vcf.gz").select("chrom", "start", "end", "ref", "alt")
|
|
354
|
-
```
|
|
355
|
-
|
|
356
|
-
5. **Use cloud paths directly:** Pass S3/GCS/Azure URIs directly to read/scan/register functions instead of downloading files first. Authenticated access uses your cloud SDK credentials (`AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, `GOOGLE_APPLICATION_CREDENTIALS`, Azure defaults) only when those cloud paths are accessed:
|
|
357
|
-
```python
|
|
358
|
-
df = pb.read_bed("s3://my-bucket/regions.bed", allow_anonymous=True)
|
|
359
|
-
```
|
|
360
|
-
|
|
361
|
-
6. **Prefer functional API for single operations, method-chaining for pipelines:** Use `pb.overlap()` for one-off operations and `.lazy().pb.overlap()` when building multi-step pipelines.
|
|
362
|
-
|
|
363
|
-
## Resources
|
|
364
|
-
|
|
365
|
-
### references/
|
|
366
|
-
|
|
367
|
-
Detailed documentation for each major capability:
|
|
368
|
-
|
|
369
|
-
- **interval_operations.md** - All 8 interval operations with parameters, examples, output schemas, and performance tips. Core reference for genomic range arithmetic.
|
|
370
|
-
|
|
371
|
-
- **file_io.md** - Supported formats table, per-format column schemas, cloud storage configuration, compression support, and common parameters.
|
|
372
|
-
|
|
373
|
-
- **sql_processing.md** - Register functions, DataFusion SQL syntax, combining SQL with interval operations, and example queries.
|
|
374
|
-
|
|
375
|
-
- **pileup_operations.md** - Per-base read depth computation from BAM/CRAM files, parameters, and integration with interval operations.
|
|
376
|
-
|
|
377
|
-
- **configuration.md** - Global settings (parallelism, coordinate systems, streaming modes), logging, and metadata management.
|
|
378
|
-
|
|
379
|
-
- **bioframe_migration.md** - Operation mapping table, API differences, performance comparison, migration code examples, and pandas compatibility mode.
|
package/skills/ponytail/SKILL.md
DELETED
|
@@ -1,31 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: ponytail
|
|
3
|
-
description: "Lazy Senior Developer Mode (Ponytail) - The best code is the code never written. Applies the 7-step engineering ladder, YAGNI, standard library reuse, and shortest working diffs."
|
|
4
|
-
risk: low
|
|
5
|
-
source: built-in
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
# Ponytail: Lazy Senior Developer Mode
|
|
9
|
-
|
|
10
|
-
You are a lazy senior developer. Lazy means efficient, not careless. The best code is the code never written.
|
|
11
|
-
|
|
12
|
-
## The 7-Step Engineering Ladder
|
|
13
|
-
|
|
14
|
-
Before writing any code, stop at the first rung that holds:
|
|
15
|
-
|
|
16
|
-
1. **YAGNI**: Does this need to be built at all? If not, do not write it.
|
|
17
|
-
2. **Codebase Reuse**: Does it already exist in this codebase? Reuse the helper, util, or pattern that is already here.
|
|
18
|
-
3. **Standard Library**: Does the runtime/stdlib already do this? Use it.
|
|
19
|
-
4. **Native Platform**: Does a native platform feature or OS tool cover it? Use it.
|
|
20
|
-
5. **Installed Dependencies**: Does an already-installed dependency solve it? Use it.
|
|
21
|
-
6. **One-Liner**: Can this be one concise, clear line? Make it one line.
|
|
22
|
-
7. **Minimum Code**: Only then write the minimum code that works.
|
|
23
|
-
|
|
24
|
-
## Core Rules
|
|
25
|
-
|
|
26
|
-
- **No unrequested abstractions**: Never introduce layers, factories, or wrappers unless explicitly requested.
|
|
27
|
-
- **No unnecessary dependencies**: Avoid adding new packages if existing tools or stdlib suffice.
|
|
28
|
-
- **No boilerplate**: Deletion over addition. Boring over clever. Fewest files possible.
|
|
29
|
-
- **Shortest working diff wins**: Fix the root cause in the shared helper once, rather than patching every caller.
|
|
30
|
-
- **Mark intentional shortcuts**: Tag intentional simplifications with `ponytail:` comments noting known ceilings and upgrade paths.
|
|
31
|
-
- **Non-trivial logic leaves a check**: Always leave one runnable check/test behind for non-trivial logic.
|
|
@@ -1,18 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: ponytail-audit
|
|
3
|
-
description: "Audit codebase for over-engineering, dead code, redundant abstractions, and unneeded dependencies."
|
|
4
|
-
risk: low
|
|
5
|
-
source: built-in
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
# Ponytail Audit: Codebase Simplification
|
|
9
|
-
|
|
10
|
-
Audit the codebase to eliminate bloat, unnecessary abstractions, and dead code.
|
|
11
|
-
|
|
12
|
-
## Audit Checklist
|
|
13
|
-
|
|
14
|
-
1. **Dead Code Scan**: Identify unused exports, dead files, and uncalled helper functions.
|
|
15
|
-
2. **Abstractions & Wrapper Layers**: Find single-use wrappers or unnecessary indirection layers that can be collapsed into straightforward direct calls.
|
|
16
|
-
3. **Redundant Dependencies**: Check `package.json` for external libraries whose functionality can be replaced with 2-3 lines of native code or stdlib.
|
|
17
|
-
4. **Boilerplate Reduction**: Consolidate duplicate patterns into single shared utilities.
|
|
18
|
-
5. **Report & Proposals**: Present candidate deletions and simplifications ordered by code reduction and maintenance savings.
|