@pikaa-ai/pikaa 0.3.23 → 0.3.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +407 -219
  6. package/dist/index.js +7 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,488 +0,0 @@
1
- ---
2
- name: umap-learn
3
- description: Use UMAP-learn for nonlinear dimensionality reduction, 2D/3D embeddings, clustering preprocessing, supervised or semi-supervised UMAP, DensMAP, AlignedUMAP, and Parametric UMAP workflows.
4
- license: BSD-3-Clause license
5
- metadata:
6
- version: "1.1"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # UMAP-Learn
11
-
12
- ## Overview
13
-
14
- UMAP (Uniform Manifold Approximation and Projection) is a dimensionality reduction technique for visualization and general non-linear dimensionality reduction. Apply this skill for fast, scalable embeddings that preserve local and global structure, supervised learning, and clustering preprocessing.
15
-
16
- ## Quick Start
17
-
18
- ### Installation
19
-
20
- Current stable release: **umap-learn 0.5.12** (released April 2026). Requires Python 3.9+ and depends on `scikit-learn>=1.6`, `numba`, `pynndescent`, `numpy`, and `scipy`. Pin to a verified release:
21
-
22
- ```bash
23
- uv pip install umap-learn==0.5.12
24
- ```
25
-
26
- ### Basic Usage
27
-
28
- UMAP follows scikit-learn conventions and can be used as a drop-in replacement for t-SNE or PCA.
29
-
30
- ```python
31
- import umap
32
- from sklearn.preprocessing import StandardScaler
33
-
34
- # Prepare data (standardization is essential)
35
- scaled_data = StandardScaler().fit_transform(data)
36
-
37
- # Method 1: Single step (fit and transform)
38
- embedding = umap.UMAP().fit_transform(scaled_data)
39
-
40
- # Method 2: Separate steps (for reusing trained model)
41
- reducer = umap.UMAP(random_state=42)
42
- reducer.fit(scaled_data)
43
- embedding = reducer.embedding_ # Access the trained embedding
44
- ```
45
-
46
- **Preprocessing requirement:** Match preprocessing to the metric. For numeric Euclidean-style metrics, scale features before fitting so high-variance columns do not dominate. For cosine, binary, precomputed-distance, or mixed-feature workflows, choose preprocessing that matches the metric instead of blindly standardizing every column.
47
-
48
- ### Typical Workflow
49
-
50
- ```python
51
- import umap
52
- import matplotlib.pyplot as plt
53
- from sklearn.preprocessing import StandardScaler
54
-
55
- # 1. Preprocess data
56
- scaler = StandardScaler()
57
- scaled_data = scaler.fit_transform(raw_data)
58
-
59
- # 2. Create and fit UMAP
60
- reducer = umap.UMAP(
61
- n_neighbors=15,
62
- min_dist=0.1,
63
- n_components=2,
64
- metric='euclidean',
65
- random_state=42
66
- )
67
- embedding = reducer.fit_transform(scaled_data)
68
-
69
- # 3. Visualize
70
- plt.scatter(embedding[:, 0], embedding[:, 1], c=labels, cmap='Spectral', s=5)
71
- plt.colorbar()
72
- plt.title('UMAP Embedding')
73
- plt.show()
74
- ```
75
-
76
- ## Parameter Tuning Guide
77
-
78
- UMAP has four primary parameters that control the embedding behavior. Understanding these is crucial for effective usage.
79
-
80
- ### n_neighbors (default: 15)
81
-
82
- **Purpose:** Balances local versus global structure in the embedding.
83
-
84
- **How it works:** Controls the size of the local neighborhood UMAP examines when learning manifold structure.
85
-
86
- **Effects by value:**
87
- - **Low values (2-5):** Emphasizes fine local detail but may fragment data into disconnected components
88
- - **Medium values (15-20):** Balanced view of both local structure and global relationships (recommended starting point)
89
- - **High values (50-200):** Prioritizes broad topological structure at the expense of fine-grained details
90
-
91
- **Recommendation:** Start with 15 and adjust based on results. Increase for more global structure, decrease for more local detail.
92
-
93
- ### min_dist (default: 0.1)
94
-
95
- **Purpose:** Controls how tightly points cluster in the low-dimensional space.
96
-
97
- **How it works:** Sets the minimum distance apart that points are allowed to be in the output representation.
98
-
99
- **Effects by value:**
100
- - **Low values (0.0-0.1):** Creates clumped embeddings useful for clustering; reveals fine topological details
101
- - **High values (0.5-0.99):** Prevents tight packing; emphasizes broad topological preservation over local structure
102
-
103
- **Recommendation:** Use 0.0 for clustering applications, 0.1-0.3 for visualization, 0.5+ for loose structure.
104
-
105
- ### n_components (default: 2)
106
-
107
- **Purpose:** Determines the dimensionality of the embedded output space.
108
-
109
- **Key feature:** Unlike t-SNE, UMAP scales well in the embedding dimension, enabling use beyond visualization.
110
-
111
- **Common uses:**
112
- - **2-3 dimensions:** Visualization
113
- - **5-10 dimensions:** Clustering preprocessing (better preserves density than 2D)
114
- - **10-50 dimensions:** Feature engineering for downstream ML models
115
-
116
- **Recommendation:** Use 2 for visualization, 5-10 for clustering, higher for ML pipelines.
117
-
118
- ### metric (default: 'euclidean')
119
-
120
- **Purpose:** Specifies how distance is calculated between input data points.
121
-
122
- **Supported metrics:**
123
- - **Minkowski variants:** euclidean, manhattan, chebyshev
124
- - **Spatial metrics:** canberra, braycurtis, haversine
125
- - **Correlation metrics:** cosine, correlation (good for text/document embeddings)
126
- - **Binary data metrics:** hamming, jaccard, dice, russellrao, kulsinski, rogerstanimoto, sokalmichener, sokalsneath, yule
127
- - **Custom metrics:** User-defined distance functions via Numba
128
-
129
- **Recommendation:** Use euclidean for numeric data, cosine for text/document vectors, hamming for binary data.
130
-
131
- ### Parameter Tuning Example
132
-
133
- ```python
134
- # For visualization with emphasis on local structure
135
- umap.UMAP(n_neighbors=15, min_dist=0.1, n_components=2, metric='euclidean')
136
-
137
- # For clustering preprocessing
138
- umap.UMAP(n_neighbors=30, min_dist=0.0, n_components=10, metric='euclidean')
139
-
140
- # For document embeddings
141
- umap.UMAP(n_neighbors=15, min_dist=0.1, n_components=2, metric='cosine')
142
-
143
- # For preserving global structure
144
- umap.UMAP(n_neighbors=100, min_dist=0.5, n_components=2, metric='euclidean')
145
- ```
146
-
147
- ## Supervised and Semi-Supervised Dimension Reduction
148
-
149
- UMAP supports incorporating label information to guide the embedding process, enabling class separation while preserving internal structure.
150
-
151
- ### Supervised UMAP
152
-
153
- Pass target labels via the `y` parameter when fitting:
154
-
155
- ```python
156
- # Supervised dimension reduction
157
- embedding = umap.UMAP().fit_transform(data, y=labels)
158
- ```
159
-
160
- **Key benefits:**
161
- - Achieves cleanly separated classes
162
- - Preserves internal structure within each class
163
- - Maintains global relationships between classes
164
-
165
- ### Semi-Supervised UMAP
166
-
167
- For partial labels, mark unlabeled points with `-1` following scikit-learn convention:
168
-
169
- ```python
170
- # Create semi-supervised labels
171
- semi_labels = labels.copy()
172
- semi_labels[unlabeled_indices] = -1
173
-
174
- # Fit with partial labels
175
- embedding = umap.UMAP().fit_transform(data, y=semi_labels)
176
- ```
177
-
178
- **When to use:** When labeling is expensive or you have more data than labels available.
179
-
180
- ## UMAP for Clustering
181
-
182
- UMAP serves as effective preprocessing for density-based clustering algorithms like HDBSCAN, overcoming the curse of dimensionality.
183
-
184
- ### Best Practices for Clustering
185
-
186
- **Key principle:** Configure UMAP differently for clustering than for visualization.
187
-
188
- **Recommended parameters:**
189
- - **n_neighbors:** Increase to ~30 (default 15 is too local and can create artificial fine-grained clusters)
190
- - **min_dist:** Set to 0.0 (pack points densely within clusters for clearer boundaries)
191
- - **n_components:** Use 5-10 dimensions (maintains performance while improving density preservation vs. 2D)
192
-
193
- ### Clustering Workflow
194
-
195
- Install HDBSCAN separately for density-based clustering:
196
-
197
- ```bash
198
- uv pip install hdbscan
199
- ```
200
-
201
- ```python
202
- import umap
203
- import hdbscan
204
- from sklearn.preprocessing import StandardScaler
205
-
206
- # 1. Preprocess data
207
- scaled_data = StandardScaler().fit_transform(data)
208
-
209
- # 2. UMAP with clustering-optimized parameters
210
- reducer = umap.UMAP(
211
- n_neighbors=30,
212
- min_dist=0.0,
213
- n_components=10, # Higher than 2 for better density preservation
214
- metric='euclidean',
215
- random_state=42
216
- )
217
- embedding = reducer.fit_transform(scaled_data)
218
-
219
- # 3. Apply HDBSCAN clustering
220
- clusterer = hdbscan.HDBSCAN(
221
- min_cluster_size=15,
222
- min_samples=5,
223
- metric='euclidean'
224
- )
225
- labels = clusterer.fit_predict(embedding)
226
-
227
- # 4. Evaluate
228
- from sklearn.metrics import adjusted_rand_score
229
- score = adjusted_rand_score(true_labels, labels)
230
- print(f"Adjusted Rand Score: {score:.3f}")
231
- print(f"Number of clusters: {len(set(labels)) - (1 if -1 in labels else 0)}")
232
- print(f"Noise points: {sum(labels == -1)}")
233
- ```
234
-
235
- ### Visualization After Clustering
236
-
237
- ```python
238
- # Create 2D embedding for visualization (separate from clustering)
239
- vis_reducer = umap.UMAP(n_neighbors=15, min_dist=0.1, n_components=2, random_state=42)
240
- vis_embedding = vis_reducer.fit_transform(scaled_data)
241
-
242
- # Plot with cluster labels
243
- import matplotlib.pyplot as plt
244
- plt.scatter(vis_embedding[:, 0], vis_embedding[:, 1], c=labels, cmap='Spectral', s=5)
245
- plt.colorbar()
246
- plt.title('UMAP Visualization with HDBSCAN Clusters')
247
- plt.show()
248
- ```
249
-
250
- **Important caveat:** UMAP does not completely preserve density and can create artificial cluster divisions. Always validate and explore resulting clusters.
251
-
252
- ## Transforming New Data
253
-
254
- UMAP enables preprocessing of new data through its `transform()` method, allowing trained models to project unseen data into the learned embedding space.
255
-
256
- ### Basic Transform Usage
257
-
258
- ```python
259
- # Train on training data
260
- trans = umap.UMAP(n_neighbors=15, random_state=42).fit(X_train)
261
-
262
- # Transform test data
263
- test_embedding = trans.transform(X_test)
264
- ```
265
-
266
- ### Integration with Machine Learning Pipelines
267
-
268
- ```python
269
- from sklearn.svm import SVC
270
- from sklearn.model_selection import train_test_split
271
- from sklearn.preprocessing import StandardScaler
272
- import umap
273
-
274
- # Split data
275
- X_train, X_test, y_train, y_test = train_test_split(data, labels, test_size=0.2)
276
-
277
- # Preprocess
278
- scaler = StandardScaler()
279
- X_train_scaled = scaler.fit_transform(X_train)
280
- X_test_scaled = scaler.transform(X_test)
281
-
282
- # Train UMAP
283
- reducer = umap.UMAP(n_components=10, random_state=42)
284
- X_train_embedded = reducer.fit_transform(X_train_scaled)
285
- X_test_embedded = reducer.transform(X_test_scaled)
286
-
287
- # Train classifier on embeddings
288
- clf = SVC()
289
- clf.fit(X_train_embedded, y_train)
290
- accuracy = clf.score(X_test_embedded, y_test)
291
- print(f"Test accuracy: {accuracy:.3f}")
292
- ```
293
-
294
- ### Important Considerations
295
-
296
- **Data consistency:** The transform method assumes the overall distribution in the higher-dimensional space is consistent between training and test data. When this assumption fails, consider using Parametric UMAP instead.
297
-
298
- **Performance:** Transform operations are efficient (typically <1 second), though initial calls may be slower due to Numba JIT compilation.
299
-
300
- **Scikit-learn compatibility:** UMAP follows standard sklearn conventions and works in pipelines. Recent 0.5.x releases also improved feature-name support and compatibility with current scikit-learn validation APIs:
301
-
302
- ```python
303
- from sklearn.pipeline import Pipeline
304
-
305
- pipeline = Pipeline([
306
- ('scaler', StandardScaler()),
307
- ('umap', umap.UMAP(n_components=10)),
308
- ('classifier', SVC())
309
- ])
310
-
311
- pipeline.fit(X_train, y_train)
312
- predictions = pipeline.predict(X_test)
313
- feature_names = pipeline.named_steps['umap'].get_feature_names_out()
314
- ```
315
-
316
- ## Advanced Features
317
-
318
- ### Parametric UMAP
319
-
320
- Parametric UMAP replaces direct embedding optimization with a learned neural network mapping function.
321
-
322
- **Key differences from standard UMAP:**
323
- - Uses TensorFlow/Keras to train encoder networks
324
- - Enables efficient transformation of new data
325
- - Supports reconstruction via decoder networks (inverse transform)
326
- - Allows custom architectures (CNNs for images, RNNs for sequences)
327
-
328
- **Installation:**
329
- ```bash
330
- uv pip install "umap-learn[parametric-umap]==0.5.12"
331
- # Installs the TensorFlow-backed Parametric UMAP extra.
332
- ```
333
-
334
- **Basic usage:**
335
- ```python
336
- from umap.parametric_umap import ParametricUMAP
337
-
338
- # Default architecture (3-layer 100-neuron fully-connected network)
339
- embedder = ParametricUMAP()
340
- embedding = embedder.fit_transform(data)
341
-
342
- # Transform new data efficiently
343
- new_embedding = embedder.transform(new_data)
344
- ```
345
-
346
- **Custom architecture:**
347
- ```python
348
- import tensorflow as tf
349
-
350
- # Define custom encoder
351
- encoder = tf.keras.Sequential([
352
- tf.keras.layers.InputLayer(shape=(input_dim,)),
353
- tf.keras.layers.Dense(128, activation='relu'),
354
- tf.keras.layers.Dense(64, activation='relu'),
355
- tf.keras.layers.Dense(2) # Output dimension
356
- ])
357
-
358
- embedder = ParametricUMAP(encoder=encoder, dims=(input_dim,))
359
- embedding = embedder.fit_transform(data)
360
- ```
361
-
362
- **Persistence:** Save Parametric UMAP with its built-in Keras-aware methods rather than plain pickle:
363
-
364
- ```python
365
- embedder.save("parametric_umap_model", exclude_raw_data=True)
366
-
367
- from umap.parametric_umap import load_ParametricUMAP
368
- loaded = load_ParametricUMAP("parametric_umap_model")
369
- new_embedding = loaded.transform(new_data)
370
- ```
371
-
372
- Recent 0.5.12 fixes include Parametric UMAP retraining stability improvements and metric-gradient fixes, so prefer the pinned current release for neural-network workflows.
373
-
374
- **When to use Parametric UMAP:**
375
- - Need efficient transformation of new data after training
376
- - Require reconstruction capabilities (inverse transforms)
377
- - Want to combine UMAP with autoencoders
378
- - Working with complex data types (images, sequences) benefiting from specialized architectures
379
-
380
- ### Inverse Transforms
381
-
382
- Inverse transforms enable reconstruction of high-dimensional data from low-dimensional embeddings.
383
-
384
- **Basic usage:**
385
- ```python
386
- reducer = umap.UMAP()
387
- embedding = reducer.fit_transform(data)
388
-
389
- # Reconstruct high-dimensional data from embedding coordinates
390
- reconstructed = reducer.inverse_transform(embedding)
391
- ```
392
-
393
- **Important limitations:**
394
- - Computationally expensive operation
395
- - Works poorly outside the convex hull of the embedding
396
- - Accuracy decreases in regions with gaps between clusters
397
-
398
- **Example: Exploring embedding space:**
399
- ```python
400
- import numpy as np
401
-
402
- # Create grid of points in embedding space
403
- x = np.linspace(embedding[:, 0].min(), embedding[:, 0].max(), 10)
404
- y = np.linspace(embedding[:, 1].min(), embedding[:, 1].max(), 10)
405
- xx, yy = np.meshgrid(x, y)
406
- grid_points = np.c_[xx.ravel(), yy.ravel()]
407
-
408
- # Reconstruct samples from grid
409
- reconstructed_samples = reducer.inverse_transform(grid_points)
410
- ```
411
-
412
- ### AlignedUMAP
413
-
414
- For analyzing temporal or related datasets (e.g., time-series experiments, batch data):
415
-
416
- ```python
417
- from umap import AlignedUMAP
418
-
419
- # List of related datasets
420
- datasets = [day1_data, day2_data, day3_data]
421
-
422
- # Relations map matching sample indices between consecutive datasets.
423
- relations = [
424
- {day1_idx: day2_idx for day1_idx, day2_idx in matched_day1_to_day2},
425
- {day2_idx: day3_idx for day2_idx, day3_idx in matched_day2_to_day3},
426
- ]
427
-
428
- # Create aligned embeddings
429
- mapper = AlignedUMAP().fit(datasets, relations=relations)
430
- aligned_embeddings = mapper.embeddings_ # List of embeddings
431
- ```
432
-
433
- **When to use:** Comparing embeddings across related datasets while maintaining consistent coordinate systems. `relations` is required for meaningful alignment; each dictionary describes how samples in one dataset correspond to samples in the next.
434
-
435
- ## Reproducibility
436
-
437
- To ensure reproducible results, always set the `random_state` parameter:
438
-
439
- ```python
440
- reducer = umap.UMAP(random_state=42)
441
- ```
442
-
443
- UMAP uses stochastic optimization, so results will vary slightly between runs without a fixed random state.
444
-
445
- Setting `random_state` prioritizes deterministic output. Leave it unset when throughput matters more than exact repeatability, because UMAP can use more parallelism without a fixed seed.
446
-
447
- ## Common Issues and Solutions
448
-
449
- **Issue:** Disconnected components or fragmented clusters
450
- - **Solution:** Increase `n_neighbors` to emphasize more global structure
451
-
452
- **Issue:** Clusters too spread out or not well separated
453
- - **Solution:** Decrease `min_dist` to allow tighter packing
454
-
455
- **Issue:** Poor clustering results
456
- - **Solution:** Use clustering-specific parameters (n_neighbors=30, min_dist=0.0, n_components=5-10)
457
-
458
- **Issue:** Transform results differ significantly from training
459
- - **Solution:** Ensure test data distribution matches training, or use Parametric UMAP
460
-
461
- **Issue:** Slow performance on large datasets
462
- - **Solution:** Set `low_memory=True` (default), or consider dimensionality reduction with PCA first
463
-
464
- **Issue:** NaN or inf values in input data
465
- - **Solution:** Impute or drop invalid rows before fitting. Current UMAP uses scikit-learn-style finite-value checks (`ensure_all_finite`) in `fit()` and `update()`, so clean numeric input is the safest default
466
-
467
- **Issue:** All points collapsed to single cluster
468
- - **Solution:** Check data preprocessing (ensure proper scaling), increase `min_dist`
469
-
470
- **Issue:** Imports resolve to a local file instead of the real package
471
- - **Solution:** Do not keep project files named `umap.py`, `sklearn.py`, `hdbscan.py`, or `tensorflow.py` beside notebooks or scripts. Those names can shadow installed packages and break or poison examples.
472
-
473
- ## Resources
474
-
475
- ### Official documentation
476
-
477
- - [UMAP user guide](https://umap-learn.readthedocs.io/en/latest/)
478
- - [Release notes](https://umap-learn.readthedocs.io/en/latest/release_notes.html)
479
- - [PyPI package](https://pypi.org/project/umap-learn/) (current stable: 0.5.12)
480
- - [GitHub repository](https://github.com/lmcinnes/umap)
481
-
482
- ### references/
483
-
484
- Contains detailed API documentation:
485
- - `api_reference.md`: Complete UMAP class parameters and methods
486
-
487
- Load these references when detailed parameter information or advanced method usage is needed.
488
-