@pikaa-ai/pikaa 0.3.23 → 0.3.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +407 -219
  6. package/dist/index.js +7 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,379 +0,0 @@
1
- ---
2
- name: polars-bio
3
- description: High-performance genomic interval operations and bioinformatics file I/O on Polars DataFrames. Overlap, nearest, merge, coverage, complement, subtract for BED/VCF/BAM/GFF intervals. Streaming, cloud-native, faster bioframe alternative.
4
- license: Apache-2.0
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.11–3.14 and polars-bio (uv pip install). Cloud I/O uses standard AWS/GCS/Azure SDK env vars when paths use s3://, gs://, or az:// URIs.
7
- metadata:
8
- version: "1.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # polars-bio
13
-
14
- ## Overview
15
-
16
- polars-bio is a high-performance Python library for genomic interval operations and bioinformatics file I/O, built on Polars, Apache Arrow, and Apache DataFusion. It provides a familiar DataFrame-centric API for interval arithmetic (overlap, nearest, merge, coverage, complement, subtract) and reading/writing common bioinformatics formats (BED, VCF, BAM, CRAM, GFF/GTF, FASTA, FASTQ).
17
-
18
- Key value propositions:
19
- - **6-38x faster** than bioframe on real-world genomic benchmarks
20
- - **Streaming/out-of-core** support for large genomes via DataFusion
21
- - **Cloud-native** file I/O (S3, GCS, Azure) with predicate pushdown
22
- - **Two API styles**: functional (`pb.overlap(df1, df2)`) and method-chaining (`df1.lazy().pb.overlap(df2)`)
23
- - **SQL interface** for genomic data via DataFusion SQL engine
24
-
25
- ## When to Use This Skill
26
-
27
- Use this skill when:
28
- - Performing genomic interval operations (overlap, nearest, merge, coverage, complement, subtract)
29
- - Reading/writing bioinformatics file formats (BED, VCF, BAM, CRAM, GFF/GTF, FASTA, FASTQ)
30
- - Processing large genomic datasets that don't fit in memory (streaming mode)
31
- - Running SQL queries on genomic data files
32
- - Migrating from bioframe to a faster alternative
33
- - Computing read depth/pileup from BAM/CRAM files
34
- - Working with Polars DataFrames containing genomic intervals
35
-
36
- ## Quick Start
37
-
38
- ### Installation
39
-
40
- Requires Python 3.11–3.14 (see [PyPI](https://pypi.org/project/polars-bio/)).
41
-
42
- ```bash
43
- uv pip install "polars-bio==0.31.0"
44
- ```
45
-
46
- For pandas compatibility (pandas ≥3.0):
47
-
48
- ```bash
49
- uv pip install "polars-bio[pandas]==0.31.0"
50
- ```
51
-
52
- ### Basic Overlap Example
53
-
54
- ```python
55
- import polars as pl
56
- import polars_bio as pb
57
-
58
- # Create two interval DataFrames
59
- df1 = pl.DataFrame({
60
- "chrom": ["chr1", "chr1", "chr1"],
61
- "start": [1, 5, 22],
62
- "end": [6, 9, 30],
63
- })
64
-
65
- df2 = pl.DataFrame({
66
- "chrom": ["chr1", "chr1"],
67
- "start": [3, 25],
68
- "end": [8, 28],
69
- })
70
-
71
- # Functional API (returns LazyFrame by default)
72
- result = pb.overlap(df1, df2)
73
- result_df = result.collect()
74
-
75
- # Get a DataFrame directly
76
- result_df = pb.overlap(df1, df2, output_type="polars.DataFrame")
77
-
78
- # Method-chaining API (via .pb accessor on LazyFrame)
79
- result = df1.lazy().pb.overlap(df2)
80
- result_df = result.collect()
81
- ```
82
-
83
- ### Reading a BED File
84
-
85
- ```python
86
- import polars_bio as pb
87
-
88
- # Eager read (loads entire file)
89
- df = pb.read_bed("regions.bed")
90
-
91
- # Lazy scan (streaming, for large files)
92
- lf = pb.scan_bed("regions.bed")
93
- result = lf.collect()
94
- ```
95
-
96
- ## Core Capabilities
97
-
98
- ### 1. Genomic Interval Operations
99
-
100
- polars-bio provides 8 core interval operations for genomic range arithmetic. All operations accept Polars DataFrames with `chrom`, `start`, `end` columns (configurable). All operations return a `LazyFrame` by default (use `output_type="polars.DataFrame"` for eager results).
101
-
102
- **Operations:**
103
- - `overlap` / `count_overlaps` - Find or count overlapping intervals between two sets (`overlap_output="left"` returns df1-only hits since 0.30.0)
104
- - `nearest` - Find nearest intervals (with configurable `k`, `overlap`, `distance` params)
105
- - `merge` - Merge overlapping/bookended intervals within a set
106
- - `cluster` - Assign cluster IDs to overlapping intervals
107
- - `coverage` - Compute per-interval coverage counts (two-input operation)
108
- - `complement` - Find gaps between intervals within a genome
109
- - `subtract` - Remove portions of intervals that overlap another set
110
-
111
- **Example:**
112
- ```python
113
- import polars_bio as pb
114
-
115
- # Find overlapping intervals (returns LazyFrame)
116
- result = pb.overlap(df1, df2, suffixes=("_1", "_2"))
117
-
118
- # Count overlaps per interval
119
- counts = pb.count_overlaps(df1, df2)
120
-
121
- # Merge overlapping intervals
122
- merged = pb.merge(df1)
123
-
124
- # Find nearest intervals
125
- nearest = pb.nearest(df1, df2)
126
-
127
- # Collect any LazyFrame result to DataFrame
128
- result_df = result.collect()
129
- ```
130
-
131
- **Reference:** See `references/interval_operations.md` for detailed documentation on all operations, parameters, output schemas, and performance considerations.
132
-
133
- ### 2. Bioinformatics File I/O
134
-
135
- Read and write common bioinformatics formats with `read_*`, `scan_*`, `write_*`, and `sink_*` functions. Supports cloud storage (S3, GCS, Azure) and compression (GZIP, BGZF).
136
-
137
- **Supported formats:**
138
- - **BED** - Genomic intervals (`read_bed`, `scan_bed`, `write_*` via generic)
139
- - **VCF** - Genetic variants (`read_vcf`, `scan_vcf`, `write_vcf`, `sink_vcf`)
140
- - **VCF Zarr** - Analysis-ready Zarr stores (`read_vcf_zarr`, `scan_vcf_zarr`; local directory paths)
141
- - **BAM** - Aligned reads (`read_bam`, `scan_bam`, `write_bam`, `sink_bam`)
142
- - **CRAM** - Compressed alignments (`read_cram`, `scan_cram`, `write_cram`, `sink_cram`)
143
- - **GFF** - Gene annotations (`read_gff`, `scan_gff`)
144
- - **GTF** - Gene annotations (`read_gtf`, `scan_gtf`)
145
- - **FASTA** - Reference sequences (`read_fasta`, `scan_fasta`, `write_fasta`, `sink_fasta`)
146
- - **FASTQ** - Sequencing reads (`read_fastq`, `scan_fastq`, `write_fastq`, `sink_fastq`)
147
- - **SAM** - Text alignments (`read_sam`, `scan_sam`, `write_sam`, `sink_sam`)
148
- - **Hi-C pairs** - Chromatin contacts (`read_pairs`, `scan_pairs`)
149
-
150
- **Example:**
151
- ```python
152
- import polars_bio as pb
153
-
154
- # Read VCF file
155
- variants = pb.read_vcf("samples.vcf.gz")
156
-
157
- # Lazy scan BAM file (streaming)
158
- alignments = pb.scan_bam("aligned.bam")
159
-
160
- # Read GFF annotations
161
- genes = pb.read_gff("annotations.gff3")
162
-
163
- # Cloud storage (individual params, not a dict)
164
- df = pb.read_bed("s3://bucket/regions.bed",
165
- allow_anonymous=True)
166
- ```
167
-
168
- **Reference:** See `references/file_io.md` for per-format column schemas, parameters, cloud storage options, and compression support.
169
-
170
- ### 3. SQL Data Processing
171
-
172
- Register bioinformatics files as tables and query them using DataFusion SQL. Combines the power of SQL with polars-bio's genomic-aware readers.
173
-
174
- ```python
175
- import polars as pl
176
- import polars_bio as pb
177
-
178
- # Register files as SQL tables (path first, name= keyword)
179
- pb.register_vcf("samples.vcf.gz", name="variants")
180
- pb.register_bed("target_regions.bed", name="regions")
181
-
182
- # Query with SQL (returns LazyFrame)
183
- result = pb.sql("SELECT chrom, start, end, ref, alt FROM variants WHERE qual > 30")
184
- result_df = result.collect()
185
-
186
- # Register a Polars DataFrame as a SQL table
187
- pb.from_polars("my_intervals", df)
188
- result = pb.sql("SELECT * FROM my_intervals WHERE chrom = 'chr1'").collect()
189
- ```
190
-
191
- **Reference:** See `references/sql_processing.md` for register functions, SQL syntax, and examples.
192
-
193
- ### 4. Pileup Operations
194
-
195
- Compute per-base read depth from BAM/CRAM files with CIGAR-aware depth calculation.
196
-
197
- ```python
198
- import polars_bio as pb
199
-
200
- # Compute depth across a BAM file
201
- depth_lf = pb.depth("aligned.bam")
202
- depth_df = depth_lf.collect()
203
-
204
- # With quality filter
205
- depth_lf = pb.depth("aligned.bam", min_mapping_quality=20)
206
- ```
207
-
208
- **Reference:** See `references/pileup_operations.md` for parameters and integration patterns.
209
-
210
- ## Key Concepts
211
-
212
- ### Coordinate Systems
213
-
214
- polars-bio defaults to **1-based** coordinates (genomic convention). This can be changed globally:
215
-
216
- ```python
217
- import polars_bio as pb
218
-
219
- # Switch to 0-based half-open coordinates (default is 1-based / False)
220
- pb.set_option("datafusion.bio.coordinate_system_zero_based", True)
221
-
222
- # Switch back to 1-based (default)
223
- pb.set_option("datafusion.bio.coordinate_system_zero_based", False)
224
- ```
225
-
226
- I/O functions also accept `use_zero_based` to set coordinate metadata on the resulting DataFrame:
227
-
228
- ```python
229
- # Read BED with explicit 0-based metadata
230
- df = pb.read_bed("regions.bed", use_zero_based=True)
231
- ```
232
-
233
- **Important:** BED files are always 0-based half-open in the file format. polars-bio handles the conversion automatically when reading BED files. Coordinate metadata is attached to DataFrames by I/O functions and propagated through operations.
234
-
235
- ### Two API Styles
236
-
237
- **Functional API** - standalone functions, explicit inputs:
238
- ```python
239
- result = pb.overlap(df1, df2, suffixes=("_1", "_2"))
240
- merged = pb.merge(df)
241
- ```
242
-
243
- **Method-chaining API** - via `.pb` accessor on **LazyFrames** (not DataFrames):
244
- ```python
245
- result = df1.lazy().pb.overlap(df2)
246
- merged = df.lazy().pb.merge()
247
- ```
248
-
249
- **Important:** The `.pb` accessor for interval operations is only available on `LazyFrame`. On `DataFrame`, `.pb` provides write operations only (`write_bam`, `write_vcf`, etc.).
250
-
251
- Method-chaining enables fluent pipelines:
252
- ```python
253
- # Chain interval operations (note: overlap outputs suffixed columns,
254
- # so rename before merge which expects chrom/start/end)
255
- result = (
256
- df1.lazy()
257
- .pb.overlap(df2)
258
- .filter(pl.col("start_2") > 1000)
259
- .select(
260
- pl.col("chrom_1").alias("chrom"),
261
- pl.col("start_1").alias("start"),
262
- pl.col("end_1").alias("end"),
263
- )
264
- .pb.merge()
265
- .collect()
266
- )
267
- ```
268
-
269
- ### Probe-Build Architecture
270
-
271
- For two-input operations (overlap, nearest, count_overlaps, coverage), polars-bio uses a probe-build join strategy:
272
- - The **first** DataFrame is the **probe** (iterated over)
273
- - The **second** DataFrame is the **build** (indexed for lookup)
274
-
275
- For best performance, pass the larger DataFrame as the first argument (probe) and the smaller one as the second (build).
276
-
277
- ### Column Conventions
278
-
279
- By default, polars-bio expects columns named `chrom`, `start`, `end`. Custom column names can be specified via lists:
280
-
281
- ```python
282
- result = pb.overlap(
283
- df1, df2,
284
- cols1=["chromosome", "begin", "finish"],
285
- cols2=["chr", "pos_start", "pos_end"],
286
- )
287
- ```
288
-
289
- ### Return Types and Collecting Results
290
-
291
- All interval operations and `pb.sql()` return a **LazyFrame** by default. Use `.collect()` to materialize results, or pass `output_type="polars.DataFrame"` for eager evaluation:
292
-
293
- ```python
294
- # Lazy (default) - collect when needed
295
- result_lf = pb.overlap(df1, df2)
296
- result_df = result_lf.collect()
297
-
298
- # Eager - get DataFrame directly
299
- result_df = pb.overlap(df1, df2, output_type="polars.DataFrame")
300
- ```
301
-
302
- ### Streaming and Out-of-Core Processing
303
-
304
- For datasets larger than available RAM, use `scan_*` functions and streaming execution:
305
-
306
- ```python
307
- # Scan files lazily
308
- lf = pb.scan_bed("large_intervals.bed")
309
-
310
- # Process with Polars streaming (requires polars ≥1.37, bundled with polars-bio)
311
- result = lf.collect(engine="streaming")
312
- ```
313
-
314
- DataFusion streaming is enabled by default for interval operations, processing data in batches without loading the full dataset into memory.
315
-
316
- ## Common Pitfalls
317
-
318
- 1. **`.pb` accessor on DataFrame vs LazyFrame:** Interval operations (overlap, merge, etc.) are only on `LazyFrame.pb`. `DataFrame.pb` only has write methods. Use `.lazy()` to convert before chaining interval ops.
319
-
320
- 2. **LazyFrame returns:** All interval operations and `pb.sql()` return `LazyFrame` by default. Don't forget `.collect()` or use `output_type="polars.DataFrame"`.
321
-
322
- 3. **Column name mismatches:** polars-bio expects `chrom`, `start`, `end` by default. Use `cols1`/`cols2` parameters (as lists) if your columns have different names.
323
-
324
- 4. **Coordinate system metadata:** Interval operations read coordinate metadata from I/O functions or DataFrame `config_meta`. For manually built DataFrames, set `df.config_meta.set(coordinate_system_zero_based=True)` (0-based) or `False` (1-based). If metadata is missing, polars-bio falls back to the global `datafusion.bio.coordinate_system_zero_based` setting (with a warning). Set `pb.set_option("datafusion.bio.coordinate_system_check", True)` to raise `MissingCoordinateSystemError` instead. Mismatched systems between inputs raise `CoordinateSystemMismatchError`.
325
-
326
- 5. **Probe-build order matters:** For overlap, nearest, and coverage, the first DataFrame is probed against the second. Swapping arguments changes which intervals appear in the left vs right output columns, and can affect performance.
327
-
328
- 6. **INT32 position limit:** Genomic positions are stored as 32-bit integers, limiting coordinates to ~2.1 billion. This is sufficient for all known genomes but may be an issue with custom coordinate spaces.
329
-
330
- 7. **BAM index requirements:** `read_bam` and `scan_bam` require a `.bai` index file alongside the BAM. Create one with `samtools index` if missing.
331
-
332
- 8. **Parallel execution disabled by default:** DataFusion parallelism defaults to 1 partition. Enable for large datasets:
333
- ```python
334
- pb.set_option("datafusion.execution.target_partitions", 8)
335
- ```
336
-
337
- 9. **CRAM has separate functions:** Use `read_cram`/`scan_cram`/`register_cram` for CRAM files (not `read_bam`). CRAM functions require a `reference_path` parameter.
338
-
339
- ## Best Practices
340
-
341
- 1. **Use `scan_*` for large files:** Prefer `scan_bed`, `scan_vcf`, etc. over `read_*` for files larger than available RAM. Scan functions enable streaming and predicate pushdown.
342
-
343
- 2. **Configure parallelism for large datasets:**
344
- ```python
345
- import os
346
- pb.set_option("datafusion.execution.target_partitions", os.cpu_count())
347
- ```
348
-
349
- 3. **Use BGZF compression:** BGZF-compressed files (`.bed.gz`, `.vcf.gz`) support parallel block decompression, significantly faster than plain GZIP.
350
-
351
- 4. **Select columns early:** When only specific columns are needed, select them early to reduce memory usage:
352
- ```python
353
- df = pb.read_vcf("large.vcf.gz").select("chrom", "start", "end", "ref", "alt")
354
- ```
355
-
356
- 5. **Use cloud paths directly:** Pass S3/GCS/Azure URIs directly to read/scan/register functions instead of downloading files first. Authenticated access uses your cloud SDK credentials (`AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, `GOOGLE_APPLICATION_CREDENTIALS`, Azure defaults) only when those cloud paths are accessed:
357
- ```python
358
- df = pb.read_bed("s3://my-bucket/regions.bed", allow_anonymous=True)
359
- ```
360
-
361
- 6. **Prefer functional API for single operations, method-chaining for pipelines:** Use `pb.overlap()` for one-off operations and `.lazy().pb.overlap()` when building multi-step pipelines.
362
-
363
- ## Resources
364
-
365
- ### references/
366
-
367
- Detailed documentation for each major capability:
368
-
369
- - **interval_operations.md** - All 8 interval operations with parameters, examples, output schemas, and performance tips. Core reference for genomic range arithmetic.
370
-
371
- - **file_io.md** - Supported formats table, per-format column schemas, cloud storage configuration, compression support, and common parameters.
372
-
373
- - **sql_processing.md** - Register functions, DataFusion SQL syntax, combining SQL with interval operations, and example queries.
374
-
375
- - **pileup_operations.md** - Per-base read depth computation from BAM/CRAM files, parameters, and integration with interval operations.
376
-
377
- - **configuration.md** - Global settings (parallelism, coordinate systems, streaming modes), logging, and metadata management.
378
-
379
- - **bioframe_migration.md** - Operation mapping table, API differences, performance comparison, migration code examples, and pandas compatibility mode.
@@ -1,31 +0,0 @@
1
- ---
2
- name: ponytail
3
- description: "Lazy Senior Developer Mode (Ponytail) - The best code is the code never written. Applies the 7-step engineering ladder, YAGNI, standard library reuse, and shortest working diffs."
4
- risk: low
5
- source: built-in
6
- ---
7
-
8
- # Ponytail: Lazy Senior Developer Mode
9
-
10
- You are a lazy senior developer. Lazy means efficient, not careless. The best code is the code never written.
11
-
12
- ## The 7-Step Engineering Ladder
13
-
14
- Before writing any code, stop at the first rung that holds:
15
-
16
- 1. **YAGNI**: Does this need to be built at all? If not, do not write it.
17
- 2. **Codebase Reuse**: Does it already exist in this codebase? Reuse the helper, util, or pattern that is already here.
18
- 3. **Standard Library**: Does the runtime/stdlib already do this? Use it.
19
- 4. **Native Platform**: Does a native platform feature or OS tool cover it? Use it.
20
- 5. **Installed Dependencies**: Does an already-installed dependency solve it? Use it.
21
- 6. **One-Liner**: Can this be one concise, clear line? Make it one line.
22
- 7. **Minimum Code**: Only then write the minimum code that works.
23
-
24
- ## Core Rules
25
-
26
- - **No unrequested abstractions**: Never introduce layers, factories, or wrappers unless explicitly requested.
27
- - **No unnecessary dependencies**: Avoid adding new packages if existing tools or stdlib suffice.
28
- - **No boilerplate**: Deletion over addition. Boring over clever. Fewest files possible.
29
- - **Shortest working diff wins**: Fix the root cause in the shared helper once, rather than patching every caller.
30
- - **Mark intentional shortcuts**: Tag intentional simplifications with `ponytail:` comments noting known ceilings and upgrade paths.
31
- - **Non-trivial logic leaves a check**: Always leave one runnable check/test behind for non-trivial logic.
@@ -1,18 +0,0 @@
1
- ---
2
- name: ponytail-audit
3
- description: "Audit codebase for over-engineering, dead code, redundant abstractions, and unneeded dependencies."
4
- risk: low
5
- source: built-in
6
- ---
7
-
8
- # Ponytail Audit: Codebase Simplification
9
-
10
- Audit the codebase to eliminate bloat, unnecessary abstractions, and dead code.
11
-
12
- ## Audit Checklist
13
-
14
- 1. **Dead Code Scan**: Identify unused exports, dead files, and uncalled helper functions.
15
- 2. **Abstractions & Wrapper Layers**: Find single-use wrappers or unnecessary indirection layers that can be collapsed into straightforward direct calls.
16
- 3. **Redundant Dependencies**: Check `package.json` for external libraries whose functionality can be replaced with 2-3 lines of native code or stdlib.
17
- 4. **Boilerplate Reduction**: Consolidate duplicate patterns into single shared utilities.
18
- 5. **Report & Proposals**: Present candidate deletions and simplifications ordered by code reduction and maintenance savings.