@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,472 +0,0 @@
1
- ---
2
- name: biopython
3
- description: Comprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use bioservices.
4
- allowed-tools: Read Write Edit Bash
5
- compatibility: Requires Python 3.10+, NumPy, and Biopython. Entrez and web BLAST examples require network access; local BLAST/MUSCLE examples require those command-line tools installed separately.
6
- license: Biopython License Agreement
7
- metadata:
8
- version: "1.2"
9
- skill-author: K-Dense Inc.
10
- openclaw:
11
- envVars:
12
- - name: NCBI_EMAIL
13
- required: false
14
- description: Email for NCBI Entrez identification (required by NCBI policy for Entrez calls).
15
- - name: NCBI_API_KEY
16
- required: false
17
- description: NCBI API key to raise Entrez rate limits.
18
- ---
19
-
20
- # Biopython: Computational Molecular Biology in Python
21
-
22
- ## Overview
23
-
24
- Biopython is a comprehensive set of freely available Python tools for biological computation. It provides functionality for sequence manipulation, file I/O, database access, structural bioinformatics, phylogenetics, and many other bioinformatics tasks. The current version is **Biopython 1.87** (released 30 March 2026). It supports **Python 3.10-3.14** and PyPy3.10, and requires NumPy. Biopython 1.87 also addresses **CVE-2025-68463** in `Bio.Entrez.Parser` when parsing untrusted files, so prefer 1.87+ for workflows that parse externally supplied Entrez XML.
25
-
26
- ## When to Use This Skill
27
-
28
- Use this skill when:
29
-
30
- - Working with biological sequences (DNA, RNA, or protein)
31
- - Reading, writing, or converting biological file formats (FASTA, GenBank, FASTQ, PDB, mmCIF, etc.)
32
- - Accessing NCBI databases (GenBank, PubMed, Protein, Gene, etc.) via Entrez
33
- - Running BLAST searches or parsing BLAST results
34
- - Performing sequence alignments (pairwise or multiple sequence alignments)
35
- - Analyzing protein structures from PDB files
36
- - Creating, manipulating, or visualizing phylogenetic trees
37
- - Finding sequence motifs or analyzing motif patterns
38
- - Calculating sequence statistics (GC content, molecular weight, melting temperature, etc.)
39
- - Performing structural bioinformatics tasks
40
- - Working with population genetics data
41
- - Any other computational molecular biology task
42
-
43
- ## Core Capabilities
44
-
45
- Biopython is organized into modular sub-packages, each addressing specific bioinformatics domains:
46
-
47
- 1. **Sequence Handling** - Bio.Seq and Bio.SeqIO for sequence manipulation and file I/O
48
- 2. **Alignment Analysis** - Bio.Align and Bio.AlignIO for pairwise and multiple sequence alignments
49
- 3. **Database Access** - Bio.Entrez for programmatic access to NCBI databases
50
- 4. **BLAST Operations** - Bio.Blast for running and parsing BLAST searches
51
- 5. **Structural Bioinformatics** - Bio.PDB for working with 3D protein structures
52
- 6. **Phylogenetics** - Bio.Phylo for phylogenetic tree manipulation and visualization
53
- 7. **Advanced Features** - Motifs, population genetics, sequence utilities, and more
54
-
55
- ## Installation and Setup
56
-
57
- Install the current stable Biopython release with an explicit version pin for reproducibility:
58
-
59
- ```bash
60
- uv pip install "biopython==1.87"
61
- ```
62
-
63
- For NCBI database access, always set your email address (required by NCBI). For reusable software, set a stable `Entrez.tool` value and register the tool/email with NCBI. For higher rate limits (10 req/s instead of 3 req/s), read only `NCBI_API_KEY` from the environment — do not hardcode keys or load unrelated environment variables:
64
-
65
- ```python
66
- import os
67
- from Bio import Entrez
68
-
69
- Entrez.email = "your.email@example.com" # required — use your real email
70
- Entrez.tool = "your_tool_name" # optional but recommended for reusable software
71
-
72
- # Optional: register at https://www.ncbi.nlm.nih.gov/account/settings/
73
- if api_key := os.environ.get("NCBI_API_KEY"):
74
- Entrez.api_key = api_key
75
- ```
76
-
77
- ## Using This Skill
78
-
79
- This skill provides comprehensive documentation organized by functionality area. When working on a task, consult the relevant reference documentation:
80
-
81
- ### 1. Sequence Handling (Bio.Seq & Bio.SeqIO)
82
-
83
- **Reference:** `references/sequence_io.md`
84
-
85
- Use for:
86
- - Creating and manipulating biological sequences
87
- - Reading and writing sequence files (FASTA, GenBank, FASTQ, etc.)
88
- - Converting between file formats
89
- - Extracting sequences from large files
90
- - Sequence translation, transcription, and reverse complement
91
- - Working with SeqRecord objects
92
-
93
- **Quick example:**
94
- ```python
95
- from Bio import SeqIO
96
-
97
- # Read sequences from FASTA file
98
- for record in SeqIO.parse("sequences.fasta", "fasta"):
99
- print(f"{record.id}: {len(record.seq)} bp")
100
-
101
- # Convert GenBank to FASTA
102
- SeqIO.convert("input.gb", "genbank", "output.fasta", "fasta")
103
- ```
104
-
105
- ### 2. Alignment Analysis (Bio.Align & Bio.AlignIO)
106
-
107
- **Reference:** `references/alignment.md`
108
-
109
- Use for:
110
- - Pairwise sequence alignment (global and local)
111
- - Reading and writing multiple sequence alignments
112
- - Using substitution matrices (BLOSUM, PAM)
113
- - Calculating alignment statistics
114
- - Customizing alignment parameters
115
-
116
- **Quick example:**
117
- ```python
118
- from Bio import Align
119
-
120
- # Pairwise alignment
121
- aligner = Align.PairwiseAligner()
122
- aligner.mode = 'global'
123
- alignments = aligner.align("ACCGGT", "ACGGT")
124
- print(alignments[0])
125
- ```
126
-
127
- ### 3. Database Access (Bio.Entrez)
128
-
129
- **Reference:** `references/databases.md`
130
-
131
- Use for:
132
- - Searching NCBI databases (PubMed, GenBank, Protein, Gene, etc.)
133
- - Downloading sequences and records
134
- - Fetching publication information
135
- - Finding related records across databases
136
- - Batch downloading with proper rate limiting
137
-
138
- **Quick example:**
139
- ```python
140
- from Bio import Entrez
141
- Entrez.email = "your.email@example.com"
142
-
143
- # Search PubMed
144
- handle = Entrez.esearch(db="pubmed", term="biopython", retmax=10)
145
- results = Entrez.read(handle)
146
- handle.close()
147
- print(f"Found {results['Count']} results")
148
- ```
149
-
150
- ### 4. BLAST Operations (Bio.Blast)
151
-
152
- **Reference:** `references/blast.md`
153
-
154
- Use for:
155
- - Running BLAST searches via NCBI web services
156
- - Running local BLAST searches
157
- - Parsing BLAST XML output
158
- - Filtering results by E-value or identity
159
- - Extracting hit sequences
160
-
161
- **Quick example:**
162
- ```python
163
- from Bio.Blast import NCBIWWW, NCBIXML
164
-
165
- # Run BLAST search
166
- result_handle = NCBIWWW.qblast("blastn", "nt", "ATCGATCGATCG")
167
- blast_record = NCBIXML.read(result_handle)
168
-
169
- # Display top hits
170
- for alignment in blast_record.alignments[:5]:
171
- print(f"{alignment.title}: E-value={alignment.hsps[0].expect}")
172
- ```
173
-
174
- ### 5. Structural Bioinformatics (Bio.PDB)
175
-
176
- **Reference:** `references/structure.md`
177
-
178
- Use for:
179
- - Parsing PDB and mmCIF structure files
180
- - Navigating protein structure hierarchy (SMCRA: Structure/Model/Chain/Residue/Atom)
181
- - Calculating distances, angles, and dihedrals
182
- - Secondary structure assignment (DSSP)
183
- - Structure superimposition and RMSD calculation
184
- - Extracting sequences from structures
185
-
186
- **Quick example:**
187
- ```python
188
- from Bio.PDB import PDBParser
189
-
190
- # Parse structure
191
- parser = PDBParser(QUIET=True)
192
- structure = parser.get_structure("1crn", "1crn.pdb")
193
-
194
- # Calculate distance between alpha carbons
195
- chain = structure[0]["A"]
196
- distance = chain[10]["CA"] - chain[20]["CA"]
197
- print(f"Distance: {distance:.2f} Å")
198
- ```
199
-
200
- ### 6. Phylogenetics (Bio.Phylo)
201
-
202
- **Reference:** `references/phylogenetics.md`
203
-
204
- Use for:
205
- - Reading and writing phylogenetic trees (Newick, NEXUS, phyloXML)
206
- - Building trees from distance matrices or alignments
207
- - Tree manipulation (pruning, rerooting, ladderizing)
208
- - Calculating phylogenetic distances
209
- - Creating consensus trees
210
- - Visualizing trees
211
-
212
- **Quick example:**
213
- ```python
214
- from Bio import Phylo
215
-
216
- # Read and visualize tree
217
- tree = Phylo.read("tree.nwk", "newick")
218
- Phylo.draw_ascii(tree)
219
-
220
- # Calculate distance
221
- distance = tree.distance("Species_A", "Species_B")
222
- print(f"Distance: {distance:.3f}")
223
- ```
224
-
225
- ### 7. Advanced Features
226
-
227
- **Reference:** `references/advanced.md`
228
-
229
- Use for:
230
- - **Sequence motifs** (Bio.motifs) - Finding and analyzing motif patterns
231
- - **Population genetics** (Bio.PopGen) - GenePop files, Fst calculations, Hardy-Weinberg tests
232
- - **Sequence utilities** (Bio.SeqUtils) - GC content, melting temperature, molecular weight, protein analysis
233
- - **Restriction analysis** (Bio.Restriction) - Finding restriction enzyme sites
234
- - **Clustering** (Bio.Cluster) - K-means and hierarchical clustering
235
- - **Genome diagrams** (GenomeDiagram) - Visualizing genomic features
236
-
237
- **Quick example:**
238
- ```python
239
- from Bio.SeqUtils import gc_fraction, molecular_weight
240
- from Bio.Seq import Seq
241
-
242
- seq = Seq("ATCGATCGATCG")
243
- print(f"GC content: {gc_fraction(seq):.2%}")
244
- print(f"Molecular weight: {molecular_weight(seq, seq_type='DNA'):.2f} g/mol")
245
- ```
246
-
247
- ## General Workflow Guidelines
248
-
249
- ### Reading Documentation
250
-
251
- When a user asks about a specific Biopython task:
252
-
253
- 1. **Identify the relevant module** based on the task description
254
- 2. **Read the appropriate reference file** using the Read tool
255
- 3. **Extract relevant code patterns** and adapt them to the user's specific needs
256
- 4. **Combine multiple modules** when the task requires it
257
-
258
- Example search patterns for reference files:
259
- ```bash
260
- # Find information about specific functions
261
- rg -n "SeqIO.parse" references/sequence_io.md
262
-
263
- # Find examples of specific tasks
264
- rg -n "BLAST" references/blast.md
265
-
266
- # Find information about specific concepts
267
- rg -n "alignment" references/alignment.md
268
- ```
269
-
270
- ### Writing Biopython Code
271
-
272
- Follow these principles when writing Biopython code:
273
-
274
- 1. **Import modules explicitly**
275
- ```python
276
- from Bio import SeqIO, Entrez
277
- from Bio.Seq import Seq
278
- ```
279
-
280
- 2. **Set Entrez email** when using NCBI databases; load only `NCBI_API_KEY` from the environment if present
281
- ```python
282
- import os
283
- from Bio import Entrez
284
-
285
- Entrez.email = "your.email@example.com"
286
- Entrez.tool = "your_tool_name"
287
- if api_key := os.environ.get("NCBI_API_KEY"):
288
- Entrez.api_key = api_key
289
- ```
290
-
291
- 3. **Use appropriate file formats** - Check which format best suits the task
292
- ```python
293
- # Common formats: "fasta", "genbank", "fastq", "clustal", "phylip"
294
- ```
295
-
296
- 4. **Handle files properly** - Close handles after use or use context managers
297
- ```python
298
- with open("file.fasta") as handle:
299
- records = SeqIO.parse(handle, "fasta")
300
- ```
301
-
302
- 5. **Use iterators for large files** - Avoid loading everything into memory
303
- ```python
304
- for record in SeqIO.parse("large_file.fasta", "fasta"):
305
- # Process one record at a time
306
- ```
307
-
308
- 6. **Handle errors gracefully** - Network operations and file parsing can fail
309
- ```python
310
- from urllib.error import HTTPError
311
-
312
- try:
313
- handle = Entrez.efetch(db="nucleotide", id=accession)
314
- except HTTPError as e:
315
- print(f"Error: {e}")
316
- ```
317
-
318
- ## Common Patterns
319
-
320
- ### Pattern 1: Fetch Sequence from GenBank
321
-
322
- ```python
323
- from Bio import Entrez, SeqIO
324
-
325
- Entrez.email = "your.email@example.com"
326
-
327
- # Fetch sequence
328
- handle = Entrez.efetch(db="nucleotide", id="EU490707", rettype="gb", retmode="text")
329
- record = SeqIO.read(handle, "genbank")
330
- handle.close()
331
-
332
- print(f"Description: {record.description}")
333
- print(f"Sequence length: {len(record.seq)}")
334
- ```
335
-
336
- ### Pattern 2: Sequence Analysis Pipeline
337
-
338
- ```python
339
- from Bio import SeqIO
340
- from Bio.SeqUtils import gc_fraction
341
-
342
- for record in SeqIO.parse("sequences.fasta", "fasta"):
343
- # Calculate statistics
344
- gc = gc_fraction(record.seq)
345
- length = len(record.seq)
346
-
347
- # Find ORFs, translate, etc.
348
- protein = record.seq.translate()
349
-
350
- print(f"{record.id}: {length} bp, GC={gc:.2%}")
351
- ```
352
-
353
- ### Pattern 3: BLAST and Fetch Top Hits
354
-
355
- ```python
356
- from Bio.Blast import NCBIWWW, NCBIXML
357
- from Bio import Entrez, SeqIO
358
-
359
- Entrez.email = "your.email@example.com"
360
-
361
- # Run BLAST
362
- result_handle = NCBIWWW.qblast("blastn", "nt", sequence)
363
- blast_record = NCBIXML.read(result_handle)
364
-
365
- # Get top hit accessions
366
- accessions = [aln.accession for aln in blast_record.alignments[:5]]
367
-
368
- # Fetch sequences
369
- for acc in accessions:
370
- handle = Entrez.efetch(db="nucleotide", id=acc, rettype="fasta", retmode="text")
371
- record = SeqIO.read(handle, "fasta")
372
- handle.close()
373
- print(f">{record.description}")
374
- ```
375
-
376
- ### Pattern 4: Build Phylogenetic Tree from Sequences
377
-
378
- ```python
379
- from Bio import AlignIO, Phylo
380
- from Bio.Phylo.TreeConstruction import DistanceCalculator, DistanceTreeConstructor
381
-
382
- # Read alignment
383
- alignment = AlignIO.read("alignment.fasta", "fasta")
384
-
385
- # Calculate distances
386
- calculator = DistanceCalculator("identity")
387
- dm = calculator.get_distance(alignment)
388
-
389
- # Build tree
390
- constructor = DistanceTreeConstructor()
391
- tree = constructor.nj(dm)
392
-
393
- # Visualize
394
- Phylo.draw_ascii(tree)
395
- ```
396
-
397
- ## Best Practices
398
-
399
- 1. **Always read relevant reference documentation** before writing code
400
- 2. **Use grep to search reference files** for specific functions or examples
401
- 3. **Validate file formats** before parsing
402
- 4. **Handle missing data gracefully** - Not all records have all fields
403
- 5. **Cache downloaded data** - Don't repeatedly download the same sequences
404
- 6. **Respect NCBI rate limits** - Use API keys, registered tool/email values for reusable software, and Entrez history/batching for large jobs
405
- 7. **Test with small datasets** before processing large files
406
- 8. **Keep Biopython updated** to get latest features and bug fixes
407
- 9. **Use appropriate genetic code tables** for translation
408
- 10. **Document analysis parameters** for reproducibility
409
-
410
- ## Troubleshooting Common Issues
411
-
412
- ### Issue: "No handlers could be found for logger 'Bio.Entrez'"
413
- **Solution:** This is just a warning. Set Entrez.email to suppress it.
414
-
415
- ### Issue: "HTTP Error 400" from NCBI
416
- **Solution:** Check that IDs/accessions are valid and properly formatted.
417
-
418
- ### Issue: "ValueError: EOF" when parsing files
419
- **Solution:** Verify file format matches the specified format string.
420
-
421
- ### Issue: Alignment fails with "sequences are not the same length"
422
- **Solution:** Ensure sequences are aligned before using AlignIO or MultipleSeqAlignment.
423
-
424
- ### Issue: BLAST searches are slow
425
- **Solution:** Use local BLAST for large-scale searches, or cache results.
426
-
427
- ### Issue: PDB parser warnings
428
- **Solution:** Use `PDBParser(QUIET=True)` to suppress warnings, or investigate structure quality.
429
-
430
- ### Issue: ImportError for Bio.HMM, Bio.MarkovModel, or Bio.Application
431
- **Solution:** These modules were removed in Biopython 1.86. Use [hmmlearn](https://pypi.org/project/hmmlearn/) for HMMs and the standard library `subprocess` module instead of `Bio.Application` CLI wrappers.
432
-
433
- ### Issue: PairwiseAligner returns fewer alignments after upgrading to 1.86+
434
- **Solution:** The default gap score changed from 0 to -1 in 1.86, eliminating trivial tie alignments. Set `aligner.gap_score = 0` to restore the old behavior if needed (see `references/alignment.md`).
435
-
436
- ## Additional Resources
437
-
438
- - **Official Documentation**: https://biopython.org/docs/latest/
439
- - **Tutorial**: https://biopython.org/docs/latest/Tutorial/
440
- - **Cookbook**: https://biopython.org/docs/latest/Tutorial/ (advanced examples)
441
- - **GitHub**: https://github.com/biopython/biopython
442
- - **Release notes**: https://github.com/biopython/biopython/blob/master/NEWS.rst
443
- - **Deprecated APIs**: https://github.com/biopython/biopython/blob/master/DEPRECATED.rst
444
- - **Mailing List**: biopython@biopython.org
445
-
446
- ## Quick Reference
447
-
448
- To locate information in reference files, use these search patterns:
449
-
450
- ```bash
451
- # Search for specific functions
452
- rg -n "function_name" references/*.md
453
-
454
- # Find examples of specific tasks
455
- rg -n "example" references/sequence_io.md
456
-
457
- # Find all occurrences of a module
458
- rg -n "Bio.Seq" references/*.md
459
- ```
460
-
461
- ## Summary
462
-
463
- Biopython provides comprehensive tools for computational molecular biology. When using this skill:
464
-
465
- 1. **Identify the task domain** (sequences, alignments, databases, BLAST, structures, phylogenetics, or advanced)
466
- 2. **Consult the appropriate reference file** in the `references/` directory
467
- 3. **Adapt code examples** to the specific use case
468
- 4. **Combine multiple modules** when needed for complex workflows
469
- 5. **Follow best practices** for file handling, error checking, and data management
470
-
471
- The modular reference documentation ensures detailed, searchable information for every major Biopython capability.
472
-