@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,470 +0,0 @@
1
- ---
2
- name: scikit-bio
3
- description: Biological data toolkit. Sequence analysis, alignments, phylogenetic trees, diversity metrics (alpha/beta, UniFrac), ordination (PCoA), PERMANOVA, FASTA/Newick I/O, for microbiome analysis.
4
- license: BSD-3-Clause license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.10+ and scikit-bio 0.7+ (uv pip install scikit-bio). NumPy 2.0+ is required. Optional matplotlib/seaborn/plotly for plotting; biom-format for BIOM tables; polars/anndata for table interoperability.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # scikit-bio
13
-
14
- ## Overview
15
-
16
- scikit-bio is a comprehensive Python library for working with biological data. Apply this skill for bioinformatics analyses spanning sequence manipulation, alignment, phylogenetics, microbial ecology, and multivariate statistics.
17
-
18
- ## When to Use This Skill
19
-
20
- This skill should be used when the user:
21
- - Works with biological sequences (DNA, RNA, protein)
22
- - Needs to read/write biological file formats (FASTA, FASTQ, GenBank, Newick, BIOM, etc.)
23
- - Performs sequence alignments or searches for motifs
24
- - Constructs or analyzes phylogenetic trees
25
- - Calculates diversity metrics (alpha/beta diversity, UniFrac distances)
26
- - Performs ordination analysis (PCoA, CCA, RDA)
27
- - Runs statistical tests on biological/ecological data (PERMANOVA, ANOSIM, Mantel)
28
- - Analyzes microbiome or community ecology data
29
- - Works with protein embeddings from language models
30
- - Needs to manipulate biological data tables
31
-
32
- ## Core Capabilities
33
-
34
- ### 1. Sequence Manipulation
35
-
36
- Work with biological sequences using specialized classes for DNA, RNA, and protein data.
37
-
38
- **Key operations:**
39
- - Read/write sequences from FASTA, FASTQ, GenBank, EMBL formats
40
- - Sequence slicing, concatenation, and searching
41
- - Reverse complement, transcription (DNA→RNA), and translation (RNA→protein)
42
- - Find motifs and patterns using regex
43
- - Calculate distances (Hamming, k-mer based)
44
- - Handle sequence quality scores and metadata
45
-
46
- **Common patterns:**
47
- ```python
48
- import skbio
49
-
50
- # Read sequences from file
51
- seq = skbio.DNA.read('input.fasta')
52
-
53
- # Sequence operations
54
- rc = seq.reverse_complement()
55
- rna = seq.transcribe()
56
- protein = rna.translate()
57
-
58
- # Find motifs
59
- motif_positions = seq.find_with_regex('ATG[ACGT]{3}')
60
-
61
- # Check for properties
62
- has_degens = seq.has_degenerates()
63
- seq_no_gaps = seq.degap()
64
- ```
65
-
66
- **Important notes:**
67
- - Use `DNA`, `RNA`, `Protein` classes for grammared sequences with validation
68
- - Use `Sequence` class for generic sequences without alphabet restrictions
69
- - Quality scores automatically loaded from FASTQ files into positional metadata
70
- - Metadata types: sequence-level (ID, description), positional (per-base), interval (regions/features)
71
-
72
- ### 2. Sequence Alignment
73
-
74
- Perform pairwise and multiple sequence alignments using the `pair_align` engine (introduced in scikit-bio 0.7.0), a versatile and efficient dynamic-programming aligner.
75
-
76
- **Key capabilities:**
77
- - Global, local, and semi-global alignment (free ends configurable) in one function
78
- - Convenience wrappers `pair_align_nucl` (BLASTN-like) and `pair_align_prot` (BLASTP-like)
79
- - Configurable scoring: match/mismatch tuple or named substitution matrix; linear or affine gap penalties
80
- - `PairAlignPath` results carry CIGAR strings and convert to aligned sequences
81
- - Multiple sequence alignment storage and manipulation with `TabularMSA`
82
-
83
- **Common patterns:**
84
- ```python
85
- from skbio import DNA, Protein
86
- from skbio.alignment import pair_align_nucl, pair_align_prot, pair_align, TabularMSA
87
-
88
- # Nucleotide alignment with BLASTN-like defaults
89
- seq1, seq2 = DNA('ACTACCAGATTACTTACGGATCAGG'), DNA('CGAAACTACTAGATTACGGATCTTA')
90
- aln = pair_align_nucl(seq1, seq2)
91
- aln.score # alignment score (float)
92
- path = aln.paths[0] # PairAlignPath (repr shows CIGAR)
93
- aligned_seqs = path.to_aligned((seq1, seq2)) # list of gapped strings
94
-
95
- # Build a TabularMSA from the alignment path + original sequences
96
- msa = TabularMSA.from_path_seqs(path, (seq1, seq2))
97
-
98
- # Customize the algorithm via pair_align (default mode='global')
99
- aln = pair_align(seq1, seq2, mode='local') # Smith-Waterman
100
- aln = pair_align(seq1, seq2, sub_score=(2, -3), gap_cost=(5, 2)) # affine gaps
101
- aln = pair_align(seq1, seq2, sub_score='NUC.4.4', gap_cost=3) # substitution matrix, linear gap
102
-
103
- # Protein alignment (BLASTP-like, BLOSUM62)
104
- aln = pair_align_prot(Protein('HEAGAWGHEE'), Protein('PAWHEAE'))
105
-
106
- # Read a multiple alignment from file and summarize
107
- msa = TabularMSA.read('alignment.fasta', constructor=DNA)
108
- consensus = msa.consensus()
109
- ```
110
-
111
- **Important notes:**
112
- - `pair_align` replaces the removed SSW wrapper (`local_pairwise_align_ssw`, `StripedSmithWaterman`) and the deprecated pure-Python aligners (`global_pairwise_align`, `local_pairwise_align_nucleotide`, etc.)
113
- - The result is a `PairAlignResult` that also unpacks as `score, paths, matrices` (use `keep_matrices=True` to retain the DP matrix)
114
- - `sub_score` accepts a `(match, mismatch)` tuple or a matrix name (e.g., `'NUC.4.4'`, `'BLOSUM62'`); `gap_cost` accepts a single number (linear) or `(open, extend)` tuple (affine)
115
- - Parse external CIGAR strings with `PairAlignPath.from_cigar('1I8M2D5M2I')`; score an existing alignment with `align_score(...)` and build a distance matrix from an MSA with `align_dists(...)`
116
-
117
- ### 3. Phylogenetic Trees
118
-
119
- Construct, manipulate, and analyze phylogenetic trees representing evolutionary relationships.
120
-
121
- **Key capabilities:**
122
- - Tree construction from distance matrices (UPGMA/WPGMA, Neighbor Joining, GME, BME)
123
- - Tree rearrangement with nearest neighbor interchange (`nni`)
124
- - Tree manipulation (pruning, rerooting, traversal)
125
- - Distance calculations (patristic via `cophenet`, Robinson-Foulds via `compare_rfd`)
126
- - ASCII visualization
127
- - Newick format I/O
128
-
129
- **Common patterns:**
130
- ```python
131
- from skbio import TreeNode
132
- from skbio.tree import nj, upgma, gme, bme, rf_dists
133
-
134
- # Read tree from file
135
- tree = TreeNode.read('tree.nwk')
136
-
137
- # Construct tree from distance matrix
138
- tree = nj(distance_matrix)
139
-
140
- # Tree operations
141
- subtree = tree.shear(['taxon1', 'taxon2', 'taxon3'])
142
- tips = [node for node in tree.tips()]
143
- lca = tree.lca(['taxon1', 'taxon2'])
144
-
145
- # Calculate distances
146
- patristic_dist = tree.find('taxon1').distance(tree.find('taxon2'))
147
- cophenetic_dm = tree.cophenet() # patristic distance matrix among tips
148
-
149
- # Compare two trees (Robinson-Foulds)
150
- rf_distance = tree.compare_rfd(other_tree)
151
- # Pairwise RF distances among many trees -> DistanceMatrix
152
- rf_dm = rf_dists([tree, other_tree, third_tree])
153
- ```
154
-
155
- **Important notes:**
156
- - Use `nj()` for neighbor joining (classic phylogenetic method)
157
- - Use `upgma()` for UPGMA/WPGMA (assumes molecular clock)
158
- - GME and BME are highly scalable for large trees; refine topology with `nni()`
159
- - `cophenet()` (formerly `tip_tip_distances`) returns the patristic distance matrix; `compare_rfd()` is the Robinson-Foulds method (`compare_wrfd`/`compare_cophenet` for weighted/cophenetic variants)
160
- - `lca()` is the lowest common ancestor; `lowest_common_ancestor` remains as an alias
161
- - Trees can be rooted or unrooted; some metrics require specific rooting
162
-
163
- ### 4. Diversity Analysis
164
-
165
- Calculate alpha and beta diversity metrics for microbial ecology and community analysis.
166
-
167
- **Key capabilities:**
168
- - Alpha diversity: richness (`sobs`, `observed_features`, `chao1`, `ace`), Shannon, Simpson, Hill numbers (`hill`), Faith's PD (`faith_pd`), generalized PD (`phydiv`), Pielou's evenness
169
- - Beta diversity: Bray-Curtis, Jaccard, weighted/unweighted UniFrac, Euclidean distances
170
- - Phylogenetic diversity metrics (require tree input)
171
- - Rarefaction and subsampling
172
- - Integration with ordination and statistical tests
173
-
174
- **Common patterns:**
175
- ```python
176
- from skbio.diversity import alpha_diversity, beta_diversity
177
-
178
- # Alpha diversity (phylogenetic metrics take taxa= for tip-name mapping)
179
- alpha = alpha_diversity('shannon', counts_matrix, ids=sample_ids)
180
- faith_pd = alpha_diversity('faith_pd', counts_matrix, ids=sample_ids,
181
- tree=tree, taxa=feature_ids)
182
-
183
- # Beta diversity
184
- bc_dm = beta_diversity('braycurtis', counts_matrix, ids=sample_ids)
185
- unifrac_dm = beta_diversity('unweighted_unifrac', counts_matrix,
186
- ids=sample_ids, tree=tree, taxa=feature_ids)
187
-
188
- # Get available metrics
189
- from skbio.diversity import get_alpha_diversity_metrics
190
- print(get_alpha_diversity_metrics())
191
- ```
192
-
193
- **Important notes:**
194
- - Counts must be integers representing abundances, not relative frequencies
195
- - The phylogenetic-metric argument is `taxa=` (renamed from `otu_ids` in 0.6.0; the old name is a deprecated alias); `observed_otus` is now `observed_features` (or `sobs`)
196
- - `counts_matrix` may be any table-like input (NumPy array, pandas/polars DataFrame, BIOM `Table`, or AnnData) via the dispatch system
197
- - Phylogenetic metrics (Faith's PD, UniFrac) require tree and taxa-to-tip mapping
198
- - Use `partial_beta_diversity()` for specific sample pairs, or `block_beta_diversity()` for large block-decomposed calculations
199
- - Alpha diversity returns a `pandas.Series`, beta diversity returns a `DistanceMatrix`
200
-
201
- ### 5. Ordination Methods
202
-
203
- Reduce high-dimensional biological data to visualizable lower-dimensional spaces.
204
-
205
- **Key capabilities:**
206
- - PCoA (Principal Coordinate Analysis) from distance matrices
207
- - CA (Correspondence Analysis) for contingency tables
208
- - CCA (Canonical Correspondence Analysis) with environmental constraints
209
- - RDA (Redundancy Analysis) for linear relationships
210
- - Biplot projection for feature interpretation
211
-
212
- **Common patterns:**
213
- ```python
214
- from skbio.stats.ordination import pcoa, cca
215
- import skbio
216
-
217
- # PCoA from distance matrix (limit dimensions for large matrices)
218
- pcoa_results = pcoa(distance_matrix, dimensions=3)
219
- pc1 = pcoa_results.samples['PC1']
220
- pc2 = pcoa_results.samples['PC2']
221
-
222
- # Built-in scatter plot colored by a metadata column
223
- fig = pcoa_results.plot(sample_metadata, column='bodysite')
224
-
225
- # CCA with environmental variables
226
- cca_results = cca(species_matrix, environmental_matrix)
227
-
228
- # Save/load ordination results
229
- pcoa_results.write('ordination.txt')
230
- results = skbio.OrdinationResults.read('ordination.txt')
231
- ```
232
-
233
- **Important notes:**
234
- - PCoA works with any distance/dissimilarity matrix; pass `dimensions` as an int (count) or a float in (0, 1] (fraction of cumulative variance to retain)
235
- - `OrdinationResults` exposes pandas-based attributes: `samples`, `features`, `eigvals`, `proportion_explained`, `biplot_scores`, `sample_constraints`
236
- - CCA reveals environmental drivers of community composition
237
- - `OrdinationResults.plot()` produces a matplotlib figure; results also integrate with seaborn/plotly
238
-
239
- ### 6. Statistical Testing
240
-
241
- Perform hypothesis tests specific to ecological and biological data.
242
-
243
- **Key capabilities:**
244
- - PERMANOVA: test group differences using distance matrices
245
- - ANOSIM: alternative test for group differences
246
- - PERMDISP: test homogeneity of group dispersions
247
- - Mantel test: correlation between distance matrices
248
- - Bioenv: find environmental variables correlated with distances
249
- - Differential abundance: `ancom`, `dirmult_ttest`, and `dirmult_lme` (longitudinal mixed-effects) in `skbio.stats.composition`
250
-
251
- **Common patterns:**
252
- ```python
253
- from skbio.stats.distance import permanova, anosim, mantel
254
-
255
- # Test if groups differ significantly
256
- permanova_results = permanova(distance_matrix, grouping, permutations=999)
257
- print(f"p-value: {permanova_results['p-value']}")
258
-
259
- # ANOSIM test
260
- anosim_results = anosim(distance_matrix, grouping, permutations=999)
261
-
262
- # Mantel test between two distance matrices
263
- mantel_results = mantel(dm1, dm2, method='pearson', permutations=999)
264
- print(f"Correlation: {mantel_results[0]}, p-value: {mantel_results[1]}")
265
-
266
- # Differential abundance on a feature table (raw counts recommended)
267
- from skbio.stats.composition import dirmult_ttest
268
- da = dirmult_ttest(counts_table, grouping, treatment='caseA', reference='control')
269
- ```
270
-
271
- **Important notes:**
272
- - Permutation tests provide non-parametric significance testing
273
- - Use 999+ permutations for robust p-values
274
- - PERMANOVA sensitive to dispersion differences; pair with PERMDISP
275
- - Mantel tests assess matrix correlation (e.g., geographic vs genetic distance)
276
- - Supply differential-abundance tests with raw counts, not pre-normalized proportions, to preserve magnitude information
277
-
278
- ### 7. File I/O and Format Conversion
279
-
280
- Read and write 19+ biological file formats with automatic format detection.
281
-
282
- **Supported formats:**
283
- - Sequences: FASTA, FASTQ, GenBank, EMBL, QSeq
284
- - Alignments: Clustal, PHYLIP, Stockholm
285
- - Trees: Newick
286
- - Tables: BIOM (HDF5 and JSON)
287
- - Distances: delimited square matrices
288
- - Analysis: BLAST+6/7, GFF3, Ordination results
289
- - Metadata: TSV/CSV with validation
290
-
291
- **Common patterns:**
292
- ```python
293
- import skbio
294
-
295
- # Read with automatic format detection
296
- seq = skbio.DNA.read('file.fasta', format='fasta')
297
- tree = skbio.TreeNode.read('tree.nwk')
298
-
299
- # Write to file
300
- seq.write('output.fasta', format='fasta')
301
-
302
- # Generator for large files (memory efficient)
303
- for seq in skbio.io.read('large.fasta', format='fasta', constructor=skbio.DNA):
304
- process(seq)
305
-
306
- # Convert formats
307
- seqs = list(skbio.io.read('input.fastq', format='fastq', constructor=skbio.DNA))
308
- skbio.io.write(seqs, format='fasta', into='output.fasta')
309
- ```
310
-
311
- **Important notes:**
312
- - Use generators for large files to avoid memory issues
313
- - Format can be auto-detected when `into` parameter specified
314
- - Some objects can be written to multiple formats
315
- - Support for stdin/stdout piping with `verify=False`
316
-
317
- ### 8. Distance Matrices
318
-
319
- Create and manipulate distance/dissimilarity matrices with statistical methods.
320
-
321
- **Key capabilities:**
322
- - Store symmetric (`DistanceMatrix`, hollow diagonal) or general pairwise (`PairwiseMatrix`) data
323
- - ID-based indexing and slicing
324
- - Integration with diversity, ordination, and statistical tests
325
- - Read/write delimited text format
326
-
327
- **Common patterns:**
328
- ```python
329
- from skbio import DistanceMatrix
330
- import numpy as np
331
-
332
- # Create from array
333
- data = np.array([[0, 1, 2], [1, 0, 3], [2, 3, 0]])
334
- dm = DistanceMatrix(data, ids=['A', 'B', 'C'])
335
-
336
- # Access distances
337
- dist_ab = dm['A', 'B']
338
- row_a = dm['A']
339
-
340
- # Read from file
341
- dm = DistanceMatrix.read('distances.txt')
342
-
343
- # Use in downstream analyses
344
- pcoa_results = pcoa(dm)
345
- permanova_results = permanova(dm, grouping)
346
- ```
347
-
348
- **Important notes:**
349
- - `DistanceMatrix` enforces symmetry and a zero (hollow) diagonal; it is a subclass of `SymmetricMatrix`
350
- - `PairwiseMatrix` (renamed from `DissimilarityMatrix`, which is kept as a deprecated alias) allows general/asymmetric values
351
- - IDs enable integration with metadata and biological knowledge
352
- - Compatible with pandas, numpy, and scikit-learn
353
-
354
- ### 9. Biological Tables
355
-
356
- Work with feature tables (OTU/ASV tables) common in microbiome research.
357
-
358
- **Key capabilities:**
359
- - BIOM format I/O (HDF5 and JSON) via the native `Table` class
360
- - Table dispatch system (0.7.0+): functions accept any `table_like` input — BIOM `Table`, pandas/polars DataFrame, NumPy array, or AnnData — without explicit conversion
361
- - Data augmentation techniques (`phylomix`, `mixup`, `aitchison_mixup`, `compos_cutmix`)
362
- - Sample/feature filtering and normalization
363
- - Metadata integration
364
-
365
- **Common patterns:**
366
- ```python
367
- from skbio import Table
368
- from skbio.diversity import beta_diversity
369
-
370
- # Read BIOM table
371
- table = Table.read('table.biom')
372
-
373
- # Access data
374
- sample_ids = table.ids(axis='sample')
375
- feature_ids = table.ids(axis='observation')
376
- counts = table.matrix_data
377
-
378
- # Filter
379
- filtered = table.filter(sample_ids_to_keep, axis='sample')
380
-
381
- # Pass table-like objects directly to scikit-bio drivers (dispatch system)
382
- import pandas as pd
383
- df = pd.read_table('data.tsv', index_col=0) # samples x features
384
- bdiv = beta_diversity('braycurtis', df) # no manual conversion needed
385
- ```
386
-
387
- **Important notes:**
388
- - BIOM tables are standard in QIIME 2 workflows
389
- - Rows typically represent samples, columns represent features (OTUs/ASVs)
390
- - Supports sparse and dense representations
391
- - With the dispatch system, functions return the same format as their input, or a user-specified output format
392
-
393
- ### 10. Protein Embeddings
394
-
395
- Work with protein language model embeddings for downstream analysis.
396
-
397
- **Key capabilities:**
398
- - Store embeddings from protein language models (ESM, ProtTrans, etc.)
399
- - Convert embeddings to distance matrices
400
- - Generate ordination objects for visualization
401
- - Export to numpy/pandas for ML workflows
402
-
403
- **Common patterns:**
404
- ```python
405
- from skbio.embedding import ProteinEmbedding, ProteinVector
406
-
407
- # Create embedding from array
408
- embedding = ProteinEmbedding(embedding_array, sequence_ids)
409
-
410
- # Convert to distance matrix for analysis
411
- dm = embedding.to_distances(metric='euclidean')
412
-
413
- # PCoA visualization of embedding space
414
- pcoa_results = embedding.to_ordination(metric='euclidean', method='pcoa')
415
-
416
- # Export for machine learning
417
- array = embedding.to_array()
418
- df = embedding.to_dataframe()
419
- ```
420
-
421
- **Important notes:**
422
- - Embeddings bridge protein language models with traditional bioinformatics
423
- - Compatible with scikit-bio's distance/ordination/statistics ecosystem
424
- - SequenceEmbedding and ProteinEmbedding provide specialized functionality
425
- - Useful for sequence clustering, classification, and visualization
426
-
427
- ## Best Practices
428
-
429
- ### Installation
430
- ```bash
431
- uv pip install scikit-bio
432
- ```
433
- Requires Python 3.10+ and NumPy 2.0+. Pre-compiled wheels are published for each release since 0.7.0, so most platforms install without a compiler. Conda users can instead run `conda install -c conda-forge scikit-bio`.
434
-
435
- ### Performance Considerations
436
- - Use generators for large sequence files to minimize memory usage
437
- - For massive phylogenetic trees, prefer GME or BME over NJ
438
- - Beta diversity calculations can be parallelized with `partial_beta_diversity()`
439
- - BIOM format (HDF5) more efficient than JSON for large tables
440
-
441
- ### Integration with Ecosystem
442
- - Sequences interoperate with Biopython via standard formats
443
- - Tables integrate with pandas, polars, and AnnData
444
- - Distance matrices compatible with scikit-learn
445
- - Ordination results visualizable with matplotlib/seaborn/plotly
446
- - Works seamlessly with QIIME 2 artifacts (BIOM, trees, distance matrices)
447
-
448
- ### Common Workflows
449
- 1. **Microbiome diversity analysis**: Read BIOM table → Calculate alpha/beta diversity → Ordination (PCoA) → Statistical testing (PERMANOVA)
450
- 2. **Phylogenetic analysis**: Read sequences → Align → Build distance matrix → Construct tree → Calculate phylogenetic distances
451
- 3. **Sequence processing**: Read FASTQ → Quality filter → Trim/clean → Find motifs → Translate → Write FASTA
452
- 4. **Comparative genomics**: Read sequences → Pairwise alignment → Calculate distances → Build tree → Analyze clades
453
-
454
- ## Reference Documentation
455
-
456
- For detailed API information, parameter specifications, and advanced usage examples, refer to `references/api_reference.md` which contains comprehensive documentation on:
457
- - Complete method signatures and parameters for all capabilities
458
- - Extended code examples for complex workflows
459
- - Troubleshooting common issues
460
- - Performance optimization tips
461
- - Integration patterns with other libraries
462
-
463
- ## Additional Resources
464
-
465
- - Official documentation: https://scikit.bio/docs/latest/
466
- - GitHub repository: https://github.com/scikit-bio/scikit-bio
467
- - Changelog: https://github.com/scikit-bio/scikit-bio/blob/main/CHANGELOG.md
468
- - Reference paper: "scikit-bio: a fundamental Python library for biological omic data," *Nature Methods* (2025), https://www.nature.com/articles/s41592-025-02981-z
469
- - Forum support: https://forum.qiime2.org (scikit-bio is part of QIIME 2 ecosystem)
470
-