@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,334 +0,0 @@
1
- ---
2
- name: esm
3
- description: Use when working directly with the `esm` Python SDK, ESM3 or ESMC model IDs, Forge/Biohub inference clients, or ESMFold2 folding workflows.
4
- license: MIT license
5
- metadata:
6
- version: "1.1"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # ESM: Evolutionary Scale Modeling
11
-
12
- ## Overview
13
-
14
- ESM provides protein language models for understanding, generating, and designing proteins. Use this skill for current EvolutionaryScale/Biohub workflows: ESM3 for generative design, ESMC for representation learning and embeddings, hosted Forge/Biohub inference, and ESMFold2 all-atom structure prediction.
15
-
16
- ## Core Capabilities
17
-
18
- ### 1. Protein Sequence Generation with ESM3
19
-
20
- Generate novel protein sequences with desired properties using multimodal generative modeling.
21
-
22
- **When to use:**
23
- - Designing proteins with specific functional properties
24
- - Completing partial protein sequences
25
- - Generating variants of existing proteins
26
- - Creating proteins with desired structural characteristics
27
-
28
- **Basic usage:**
29
-
30
- ```python
31
- from esm.models.esm3 import ESM3
32
- from esm.sdk.api import ESM3InferenceClient, ESMProtein, GenerationConfig
33
-
34
- # Load local open weights after accepting the license on Hugging Face.
35
- model: ESM3InferenceClient = ESM3.from_pretrained("esm3-open").to("cuda")
36
-
37
- # Create protein prompt
38
- protein = ESMProtein(sequence="MPRT___KEND") # '_' represents masked positions
39
-
40
- # Generate completion
41
- protein = model.generate(protein, GenerationConfig(track="sequence", num_steps=8))
42
- print(protein.sequence)
43
- ```
44
-
45
- **For remote/cloud usage via Forge API:**
46
-
47
- ```python
48
- import os
49
- import esm
50
- from esm.sdk.api import ESMProtein, GenerationConfig
51
-
52
- # Same interface as local ESM3; token from ESM_API_KEY (see Authentication)
53
- model = esm.sdk.client("esm3-medium-2024-08", token=os.environ["ESM_API_KEY"])
54
-
55
- # Generate
56
- protein = model.generate(protein, GenerationConfig(track="sequence", num_steps=8))
57
- ```
58
-
59
- See `references/esm3-api.md` for detailed ESM3 model specifications, advanced generation configurations, and multimodal prompting examples.
60
-
61
- ### 2. Structure Prediction and Inverse Folding
62
-
63
- Use ESM3's structure track for structure prediction from sequence or inverse folding (sequence design from structure).
64
-
65
- **Structure prediction:**
66
-
67
- ```python
68
- from esm.sdk.api import ESM3InferenceClient, ESMProtein, GenerationConfig
69
-
70
- # Predict structure from sequence
71
- protein = ESMProtein(sequence="MPRTKEINDAGLIVHSP...")
72
- protein_with_structure = model.generate(
73
- protein,
74
- GenerationConfig(track="structure", num_steps=protein.sequence.count("_"))
75
- )
76
-
77
- # Access predicted structure
78
- coordinates = protein_with_structure.coordinates # 3D coordinates
79
- pdb_string = protein_with_structure.to_pdb()
80
- ```
81
-
82
- **Inverse folding (sequence from structure):**
83
-
84
- ```python
85
- # Design sequence for a target structure
86
- protein_with_structure = ESMProtein.from_pdb("target_structure.pdb")
87
- protein_with_structure.sequence = None # Remove sequence
88
-
89
- # Generate sequence that folds to this structure
90
- designed_protein = model.generate(
91
- protein_with_structure,
92
- GenerationConfig(track="sequence", num_steps=50, temperature=0.7)
93
- )
94
- ```
95
-
96
- ### 3. Protein Embeddings with ESM C
97
-
98
- Generate high-quality embeddings for downstream tasks like function prediction, classification, or similarity analysis.
99
-
100
- **When to use:**
101
- - Extracting protein representations for machine learning
102
- - Computing sequence similarities
103
- - Feature extraction for protein classification
104
- - Transfer learning for protein-related tasks
105
-
106
- **Basic usage:**
107
-
108
- ```python
109
- from esm.models.esmc import ESMC
110
- from esm.sdk.api import ESMProtein, LogitsConfig
111
-
112
- # Load ESM C model
113
- model = ESMC.from_pretrained("esmc_300m").to("cuda")
114
-
115
- # Get embeddings
116
- protein = ESMProtein(sequence="MPRTKEINDAGLIVHSP...")
117
- protein_tensor = model.encode(protein)
118
- logits_output = model.logits(
119
- protein_tensor,
120
- LogitsConfig(sequence=True, return_embeddings=True),
121
- )
122
- embeddings = logits_output.embeddings
123
- ```
124
-
125
- **Batch processing:**
126
-
127
- ```python
128
- # Encode multiple proteins
129
- proteins = [
130
- ESMProtein(sequence="MPRTKEIND..."),
131
- ESMProtein(sequence="AGLIVHSPQ..."),
132
- ESMProtein(sequence="KTEFLNDGR...")
133
- ]
134
-
135
- embeddings_list = [
136
- model.logits(
137
- model.encode(p),
138
- LogitsConfig(sequence=True, return_embeddings=True),
139
- ).embeddings
140
- for p in proteins
141
- ]
142
- ```
143
-
144
- See `references/esm-c-api.md` for ESM C model details, efficiency comparisons, and advanced embedding strategies.
145
-
146
- ### 4. Function Conditioning and Annotation
147
-
148
- Use ESM3's function track to generate proteins with specific functional annotations or predict function from sequence.
149
-
150
- **Function-conditioned generation:**
151
-
152
- ```python
153
- from esm.sdk.api import ESMProtein, FunctionAnnotation, GenerationConfig
154
-
155
- # Create protein with desired function
156
- protein = ESMProtein(
157
- sequence="_" * 200, # Generate 200 residue protein
158
- function_annotations=[
159
- FunctionAnnotation(label="fluorescent_protein", start=50, end=150)
160
- ]
161
- )
162
-
163
- # Generate sequence with specified function
164
- functional_protein = model.generate(
165
- protein,
166
- GenerationConfig(track="sequence", num_steps=200)
167
- )
168
- ```
169
-
170
- ### 5. Chain-of-Thought Generation
171
-
172
- Iteratively refine protein designs using ESM3's chain-of-thought generation approach.
173
-
174
- ```python
175
- from esm.sdk.api import GenerationConfig
176
-
177
- # Multi-step refinement
178
- protein = ESMProtein(sequence="MPRT" + "_" * 100 + "KEND")
179
-
180
- # Step 1: Generate initial structure
181
- config = GenerationConfig(track="structure", num_steps=50)
182
- protein = model.generate(protein, config)
183
-
184
- # Step 2: Refine sequence based on structure
185
- config = GenerationConfig(track="sequence", num_steps=50, temperature=0.5)
186
- protein = model.generate(protein, config)
187
-
188
- # Step 3: Predict function
189
- config = GenerationConfig(track="function", num_steps=20)
190
- protein = model.generate(protein, config)
191
- ```
192
-
193
- ### 6. Batch Processing with Forge API
194
-
195
- Process multiple proteins efficiently using Forge's async methods.
196
-
197
- ```python
198
- import os
199
- import asyncio
200
- import esm
201
- from esm.sdk.api import ESMProtein, GenerationConfig
202
-
203
- client = esm.sdk.client("esm3-medium-2024-08", token=os.environ["ESM_API_KEY"])
204
-
205
- # Async batch processing
206
- async def batch_generate(proteins_list):
207
- tasks = [
208
- client.async_generate(protein, GenerationConfig(track="sequence"))
209
- for protein in proteins_list
210
- ]
211
- return await asyncio.gather(*tasks)
212
-
213
- # Execute
214
- proteins = [ESMProtein(sequence=f"MPRT{'_' * 50}KEND") for _ in range(10)]
215
- results = asyncio.run(batch_generate(proteins))
216
- ```
217
-
218
- See `references/forge-api.md` for detailed Forge API documentation, authentication, rate limits, and batch processing patterns.
219
-
220
- ## Model Selection Guide
221
-
222
- **ESM3 Models (Generative):**
223
- - `esm3-open` (1.4B) - Open weights, local usage after accepting the Hugging Face license
224
- - `esm3-medium-2024-08` (7B) - Best balance of quality and speed (Forge only)
225
- - `esm3-large-2024-03` (98B) - Highest quality, slower (Forge only)
226
-
227
- **ESM C Models (Embeddings):**
228
- - `esmc_300m` / `esmc-300m-2024-12` (30 layers) - Lightweight, fast inference (open weights, local)
229
- - `esmc_600m` / `esmc-600m-2024-12` (36 layers) - Balanced performance (open weights, local)
230
- - `esmc-6b-2024-12` (80 layers) - Maximum quality (Forge API; local 6B weights require Forge or SageMaker)
231
-
232
- Local `ESMC.from_pretrained()` examples use underscore aliases (`esmc_300m`, `esmc_600m`). Hosted API clients use dated model IDs such as `esmc-600m-2024-12`.
233
-
234
- **Selection criteria:**
235
- - **Local development/testing:** Use `esm3-open` or `esmc_300m`
236
- - **Production quality:** Use `esm3-medium-2024-08` via Forge
237
- - **Maximum accuracy:** Use `esm3-large-2024-03` or `esmc-6b-2024-12` via Forge
238
- - **High throughput:** Use Forge or Biohub APIs with explicit async concurrency limits
239
- - **Cost optimization:** Use smaller models, implement caching strategies
240
-
241
- ## Installation
242
-
243
- Install from PyPI ([`esm` on PyPI](https://pypi.org/project/esm/) by EvolutionaryScale). Current PyPI release: **3.2.3** (Oct 14, 2025). Requires **Python >=3.12,<3.13**.
244
-
245
- **Basic installation:**
246
-
247
- ```bash
248
- uv pip install "esm==3.2.3"
249
- ```
250
-
251
- **With Flash Attention (recommended for faster inference on NVIDIA GPUs):**
252
-
253
- ```bash
254
- uv pip install "esm==3.2.3"
255
- uv pip install flash-attn --no-build-isolation
256
- ```
257
-
258
- The Forge client ships with the `esm` package - no extra install for ESM3 or ESMC Forge inference.
259
-
260
- ## Authentication
261
-
262
- Forge API access requires an API key. Never hardcode tokens in scripts or commit them to version control.
263
-
264
- 1. Check whether `ESM_API_KEY` is already set in the environment.
265
- 2. If not, check a local `.env` for `ESM_API_KEY` only (do not load unrelated secrets).
266
- 3. If still missing, create a key in the [Biohub developer console](https://biohub.ai/developer-console/api-keys) for Biohub APIs or [Forge](https://forge.evolutionaryscale.ai) for legacy Forge-hosted ESM3/ESMC access.
267
-
268
- ```python
269
- import os
270
-
271
- token = os.environ["ESM_API_KEY"] # raises KeyError if unset
272
- ```
273
-
274
- `esm.sdk.client()` reads `ESM_API_KEY` automatically when `token` is omitted. Keep endpoint URLs fixed to trusted hosts such as `https://forge.evolutionaryscale.ai` or `https://biohub.ai`; do not take API hosts from untrusted user input.
275
-
276
- **Biohub platform:** EvolutionaryScale and Forge now surface current hosted models through [biohub.ai](https://biohub.ai). SDK class names may still reference "Forge". See `references/biohub-platform.md` for ESMFold2 and Biohub-specific setup.
277
-
278
- ## Common Workflows
279
-
280
- For detailed examples and complete workflows, see `references/workflows.md` which includes:
281
- - Novel GFP design with chain-of-thought
282
- - Protein variant generation and screening
283
- - Structure-based sequence optimization
284
- - Function prediction pipelines
285
- - Embedding-based clustering and analysis
286
-
287
- ## References
288
-
289
- This skill includes comprehensive reference documentation:
290
-
291
- - `references/esm3-api.md` - ESM3 model architecture, API reference, generation parameters, and multimodal prompting
292
- - `references/esm-c-api.md` - ESM C model details, embedding strategies, and performance optimization
293
- - `references/forge-api.md` - Forge platform documentation, authentication, batch processing, and deployment
294
- - `references/biohub-platform.md` - Biohub API migration, ESMFold2 structure prediction, and developer-console auth
295
- - `references/workflows.md` - Complete examples and common workflow patterns
296
-
297
- These references contain detailed API specifications, parameter descriptions, and advanced usage patterns. Load them as needed for specific tasks.
298
-
299
- ## Best Practices
300
-
301
- **For generation tasks:**
302
- - Start with smaller models for prototyping (`esm3-open`)
303
- - Use temperature parameter to control diversity (0.0 = deterministic, 1.0 = diverse)
304
- - Implement iterative refinement with chain-of-thought for complex designs
305
- - Validate generated sequences with structure prediction or wet-lab experiments
306
-
307
- **For embedding tasks:**
308
- - Batch process sequences when possible for efficiency
309
- - Cache embeddings for repeated analyses
310
- - Normalize embeddings when computing similarities
311
- - Use appropriate model size based on downstream task requirements
312
-
313
- **For production deployment:**
314
- - Use Forge API for scalability and latest models
315
- - Implement error handling and retry logic for API calls
316
- - Monitor token usage and implement rate limiting
317
- - Consider AWS SageMaker deployment for dedicated infrastructure
318
-
319
- ## Resources and Documentation
320
-
321
- - **GitHub Repository:** https://github.com/Biohub/esm (current ESMC/ESMFold2/Biohub docs; ESM3 docs remain linked from the repository)
322
- - **Forge Platform:** https://forge.evolutionaryscale.ai
323
- - **Biohub Platform:** https://biohub.ai
324
- - **Scientific Paper:** Hayes et al., Science (2025) - https://www.science.org/doi/10.1126/science.ads0018
325
- - **Blog Posts:**
326
- - ESM3 Release: https://www.evolutionaryscale.ai/blog/esm3-release
327
- - ESM C Launch: https://www.evolutionaryscale.ai/blog/esm-cambrian
328
- - **Community:** Slack community at https://bit.ly/3FKwcWd
329
- - **Model Weights:** Hugging Face EvolutionaryScale and Biohub organizations
330
-
331
- ## Responsible Use
332
-
333
- ESM is designed for beneficial applications in protein engineering, drug discovery, and scientific research. Follow the Responsible Biodesign Framework (https://responsiblebiodesign.ai/) and Biohub Acceptable Use Policy (https://biohub.org/acceptable-use-policy/) when designing novel proteins. Consider biosafety and ethical implications of protein designs before experimental validation.
334
-
@@ -1,327 +0,0 @@
1
- ---
2
- name: etetoolkit
3
- description: Analyze, manipulate, compare, annotate, and visualize phylogenetic or other hierarchical trees with ETE 4. Use for Newick/Nexus tree I/O, topology edits and pattern matching, Robinson-Foulds comparisons, gene-tree evolutionary events and reconciliation, NCBI/GTDB taxonomy, SmartView exploration, and publication rendering. Do not use it to infer trees from raw sequences; align sequences and infer a tree first.
4
- license: GPL-3.0-or-later
5
- allowed-tools: Read Write Edit Bash Python
6
- compatibility: Bundled scripts require Python 3.10+ and ete4 4.4.0 (upstream ete4 supports Python >=3.7). Taxonomy setup and SmartView exploration need network access; static SmartView PNG rendering needs ete4[render-sm], and Qt PDF/SVG rendering needs ete4[treeview].
7
- metadata:
8
- version: "2.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # ETE Toolkit 4
13
-
14
- ## Scope
15
-
16
- Use ETE 4 to work with an existing tree:
17
-
18
- - Read Newick/Nexus, then inspect, annotate, transform, root, prune, and write
19
- Newick trees
20
- - Compare topologies and calculate phylogenetic distances
21
- - Find repeated subtree topologies with `TreePattern`
22
- - Analyze gene trees with `PhyloTree`
23
- - Query local NCBI or GTDB taxonomy databases
24
- - Explore large trees interactively with SmartView
25
- - Render PNG with SmartView or PNG/PDF/SVG with the optional Qt treeview
26
-
27
- ETE does not replace sequence alignment or phylogenetic inference software. For
28
- raw sequences, first use MAFFT or another aligner and IQ-TREE 2, FastTree, or
29
- another inference tool; then load the resulting tree into ETE.
30
-
31
- ## Current Target
32
-
33
- This skill targets **ETE 4.4.0**, released September 3, 2025 and verified as the
34
- current PyPI release on July 23, 2026.
35
-
36
- Use `https://etetoolkit.github.io/ete/` for ETE 4 documentation. The
37
- `etetoolkit.org/docs/latest` pages are legacy ETE 3 documentation despite the
38
- URL name.
39
-
40
- Do not silently translate these examples back to ETE 3:
41
-
42
- - Package and import: `ete4`, not `ete3`
43
- - File input: pass an open file object; use strings for Newick text and do not
44
- rely on path-string heuristics retained in ETE 4.4.0
45
- - Newick selection: `parser=`, not `format=`
46
- - Node metadata: `props`, `add_prop()`, and `add_props()`
47
- - Iteration: `leaves()`, `descendants()`, and related methods return iterators
48
- - Predicates: `node.is_leaf` and `node.is_root` are properties, not methods
49
- - Node lookup: `tree["name"]`, not `tree & "name"`
50
-
51
- For porting older code, load
52
- [`references/migration-ete3-to-ete4.md`](references/migration-ete3-to-ete4.md).
53
-
54
- ## Installation
55
-
56
- Install the pinned base package:
57
-
58
- ```bash
59
- uv pip install "ete4==4.4.0"
60
- ```
61
-
62
- Add only the visualization extra required by the workflow:
63
-
64
- ```bash
65
- # SmartView static PNG screenshots
66
- uv pip install "ete4[render-sm]==4.4.0"
67
-
68
- # Legacy Qt renderer for PNG, PDF, and SVG
69
- uv pip install "ete4[treeview]==4.4.0"
70
- ```
71
-
72
- Confirm the active environment:
73
-
74
- ```bash
75
- uv run --with "ete4==4.4.0" python -c "import ete4; print(ete4.__version__)"
76
- ```
77
-
78
- No credentials are required. NCBI and GTDB workflows download public taxonomy
79
- data and can consume substantial disk space; see
80
- [`references/taxonomy.md`](references/taxonomy.md) before the first update.
81
-
82
- ## Quick Start
83
-
84
- ```python
85
- from pathlib import Path
86
-
87
- from ete4 import Tree
88
-
89
- # Use an open file object for files; reserve strings for Newick text.
90
- with Path("tree.nw").open(encoding="utf-8") as handle:
91
- tree = Tree(handle, parser=1) # parser 1: internal node names
92
-
93
- print(tree.to_str(props=["name", "dist"], compact=True))
94
- print("Leaves:", list(tree.leaf_names()))
95
-
96
- # Search and annotate.
97
- focal = tree["species1"]
98
- focal.add_props(host="human", status="focal")
99
-
100
- # Keep selected tips while preserving pairwise branch-length distances.
101
- tree.prune(
102
- ["species1", "species2", "species3"],
103
- preserve_branch_length=True,
104
- )
105
-
106
- # Root and serialize explicitly.
107
- tree.set_midpoint_outgroup()
108
- tree.write(
109
- outfile="processed.nw",
110
- parser=1,
111
- props=["host", "status"],
112
- )
113
- ```
114
-
115
- Choose the parser deliberately. A parser mismatch is the most common cause of
116
- `NewickError`, lost internal labels, or support values being read as names.
117
- See [`references/api_reference.md`](references/api_reference.md).
118
-
119
- ## Core Workflows
120
-
121
- ### Inspect and transform a tree
122
-
123
- ```python
124
- from ete4 import Tree
125
-
126
- tree = Tree("((A:1,B:1)CladeAB:0.4,C:2)Root;", parser=1)
127
-
128
- for node in tree.traverse("preorder"):
129
- label = node.name if node.name is not None else node.id
130
- print(label, node.level, node.is_leaf, node.dist)
131
-
132
- tree["A"].add_prop("group", "case")
133
- tree["B"].add_prop("group", "control")
134
-
135
- mrca = tree.common_ancestor("A", "B")
136
- print(mrca.name)
137
-
138
- tree.write(
139
- outfile="annotated.nhx",
140
- parser=1,
141
- props=["group"],
142
- format_root_node=True,
143
- )
144
- ```
145
-
146
- Node names need not be unique. `tree["A"]` returns the first match; use
147
- `list(tree.search_nodes(name="A"))` and validate the count when duplicates are
148
- possible.
149
-
150
- ### Compare two topologies
151
-
152
- ```python
153
- from ete4 import Tree
154
-
155
- tree_a = Tree("((A,B),(C,D));")
156
- tree_b = Tree("((A,C),(B,D));")
157
-
158
- (
159
- rf,
160
- max_rf,
161
- common_leaves,
162
- edges_a,
163
- edges_b,
164
- discarded_a,
165
- discarded_b,
166
- ) = tree_a.robinson_foulds(tree_b)
167
-
168
- normalized_rf = rf / max_rf if max_rf else 0.0
169
- print(rf, max_rf, normalized_rf, sorted(common_leaves))
170
- ```
171
-
172
- RF comparison uses shared leaf labels and requires meaningful, preferably
173
- unique names. Decide explicitly whether rooted or unrooted comparison is
174
- scientifically appropriate.
175
-
176
- ### Detect duplication and speciation events
177
-
178
- ```python
179
- from ete4 import PhyloTree
180
-
181
- gene_tree = PhyloTree(
182
- "((Hsa|g1,Ptr|g1),(Hsa|g2,Mmu|g1));",
183
- sp_naming_function=lambda name: name.split("|", 1)[0],
184
- )
185
-
186
- for event in gene_tree.get_descendant_evol_events(sos_thr=0.0):
187
- relationship = "speciation/orthology" if event.etype == "S" else "duplication/paralogy"
188
- print(relationship, sorted(event.in_seqs), sorted(event.out_seqs))
189
- ```
190
-
191
- Species-overlap calls are inferences from the supplied topology and naming
192
- function, not independent evidence of orthology. Pass the naming function
193
- explicitly, and use a rooted, fully bifurcating gene tree. For strict
194
- reconciliation, use a curated species tree and
195
- `gene_tree.reconcile(species_tree)`.
196
-
197
- ### Query taxonomy
198
-
199
- ```python
200
- from ete4 import NCBITaxa
201
-
202
- ncbi = NCBITaxa()
203
- names = ["Homo sapiens", "Pan troglodytes", "Mus musculus"]
204
- name_to_taxids = ncbi.get_name_translator(names)
205
-
206
- missing = [name for name in names if name not in name_to_taxids]
207
- if missing:
208
- raise ValueError(f"Names not resolved by NCBI taxonomy: {missing}")
209
-
210
- taxids = [name_to_taxids[name][0] for name in names]
211
- taxonomy_tree = ncbi.get_topology(taxids)
212
- print(taxonomy_tree.to_str(props=["sci_name", "rank"]))
213
- ```
214
-
215
- ETE 4 also provides `GTDBTaxa` for genome-centric bacterial and archaeal
216
- taxonomy. Do not mix NCBI numeric TaxIDs and GTDB string identifiers.
217
-
218
- ### Visualize
219
-
220
- Interactive SmartView:
221
-
222
- ```python
223
- from ete4 import Tree
224
-
225
- tree = Tree("((A:1,B:1)90:0.2,C:1);", parser="support")
226
- tree.explore()
227
- ```
228
-
229
- Static SmartView screenshot:
230
-
231
- ```python
232
- tree.render_sm("tree.png", w=1200, h=800)
233
- ```
234
-
235
- `render_sm()` produces PNG screenshot data; use the Qt treeview renderer when
236
- the deliverable must be vector PDF or SVG. Load
237
- [`references/visualization.md`](references/visualization.md) for layouts,
238
- faces, remote exploration, and renderer selection.
239
-
240
- ## Bundled Scripts
241
-
242
- Run from this skill directory. The commands below use a pinned, isolated ETE 4
243
- runtime through `uv run --with`.
244
-
245
- ### Tree operations
246
-
247
- ```bash
248
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
249
- stats tree.nw --parser 1
250
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
251
- ascii tree.nw --parser 1 --props name,dist
252
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
253
- convert tree.nw output.nw \
254
- --input-parser 1 --output-parser 1
255
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
256
- reroot tree.nw rooted.nw \
257
- --parser 1 --midpoint
258
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
259
- prune tree.nw pruned.nw \
260
- --parser 1 --keep species1 species2 species3
261
- uv run --with "ete4==4.4.0" python scripts/tree_operations.py \
262
- compare tree_a.nw tree_b.nw
263
- ```
264
-
265
- Use `--keep-file taxa.txt` instead of `--keep ...` for one taxon per line.
266
- The script refuses ambiguous or missing requested names rather than silently
267
- producing a partial tree.
268
-
269
- ### Visualization
270
-
271
- ```bash
272
- # Interactive SmartView
273
- uv run --with "ete4==4.4.0" python scripts/quick_visualize.py \
274
- tree.nw --parser 1
275
-
276
- # SmartView PNG (requires ete4[render-sm])
277
- uv run --with "ete4[render-sm]==4.4.0" python scripts/quick_visualize.py \
278
- tree.nw tree.png \
279
- --parser support --mode circular --show-support --color-by-support
280
-
281
- # Vector output via Qt treeview (requires ete4[treeview])
282
- uv run --with "ete4[treeview]==4.4.0" python scripts/quick_visualize.py \
283
- tree.nw tree.svg \
284
- --parser 1 --engine treeview --title "Species phylogeny"
285
- ```
286
-
287
- ## Quality and Interpretation Checks
288
-
289
- Before reporting a result:
290
-
291
- 1. Confirm the parser preserves the intended internal names, support, and
292
- branch lengths.
293
- 2. Check for empty and duplicate leaf names before name-based lookup or RF
294
- comparison.
295
- 3. State whether the tree is treated as rooted or unrooted.
296
- 4. Preserve branch lengths when pruning only if retained pairwise distances
297
- should remain unchanged.
298
- 5. Treat arbitrary polytomy resolution as a display/algorithmic convenience,
299
- not evolutionary evidence.
300
- 6. Record ETE version, parser, rooting method, pruning set, and taxonomy
301
- database snapshot in reproducible analyses.
302
- 7. Prefer iterators for large trees and `get_cached_content()` for repeated
303
- descendant-content queries.
304
-
305
- ## Reference Map
306
-
307
- Load only the reference needed for the task:
308
-
309
- - [`references/api_reference.md`](references/api_reference.md) — ETE 4 core
310
- classes, parsers, properties, traversal, I/O, topology, and comparison
311
- - [`references/workflows.md`](references/workflows.md) — complete analysis
312
- patterns, validation, reconciliation, batching, and large-tree work
313
- - [`references/visualization.md`](references/visualization.md) — SmartView,
314
- layouts/faces, PNG screenshots, and Qt vector rendering
315
- - [`references/taxonomy.md`](references/taxonomy.md) — NCBI and GTDB setup,
316
- translation, topology, annotation, and reproducibility
317
- - [`references/migration-ete3-to-ete4.md`](references/migration-ete3-to-ete4.md)
318
- — breaking API changes and porting checklist
319
-
320
- ## Authoritative Upstream Sources
321
-
322
- - Documentation: https://etetoolkit.github.io/ete/
323
- - ETE 3 to ETE 4 migration: https://etetoolkit.github.io/ete/3to4.html
324
- - Releases: https://github.com/etetoolkit/ete/releases
325
- - PyPI: https://pypi.org/project/ete4/
326
- - Source: https://github.com/etetoolkit/ete
327
- - Visualization gallery: https://github.com/etetoolkit/ete-gallery