@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,431 +0,0 @@
1
- ---
2
- name: anndata
3
- description: Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
4
- license: BSD-3-Clause license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.11+ and uv. Examples target AnnData 0.12.16, with experimental APIs clearly marked where used.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # AnnData
13
-
14
- ## Overview
15
-
16
- AnnData is a Python package for handling annotated data matrices, storing experimental measurements (X) alongside observation metadata (obs), variable metadata (var), and multi-dimensional annotations (obsm, varm, obsp, varp, uns). Originally designed for single-cell genomics through Scanpy, it now serves as a general-purpose framework for any annotated data requiring efficient storage, manipulation, and analysis.
17
-
18
- ## When to Use This Skill
19
-
20
- Use this skill when:
21
- - Creating, reading, or writing AnnData objects
22
- - Working with h5ad, zarr, or other genomics data formats
23
- - Performing single-cell RNA-seq analysis
24
- - Managing large datasets with sparse matrices or backed mode
25
- - Concatenating multiple datasets or experimental batches
26
- - Subsetting, filtering, or transforming annotated data
27
- - Integrating with scanpy, scvi-tools, or other scverse ecosystem tools
28
-
29
- ## Installation
30
-
31
- Requires Python 3.11+. Current stable release: 0.12.16 (released 2026-05-18).
32
-
33
- ```bash
34
- uv pip install "anndata==0.12.16"
35
-
36
- # Lazy I/O and dask-backed operations
37
- uv pip install "anndata[dask,lazy]==0.12.16"
38
-
39
- # Development / docs (contributors)
40
- uv pip install "anndata[dev,test,doc]==0.12.16"
41
- ```
42
-
43
- Use unpinned installs only when intentionally tracking the latest compatible release.
44
-
45
- Current API notes:
46
- - Use `anndata.io` for non-native `read_*` and `write_*` helpers. Top-level `anndata.read_h5ad` and `anndata.read_zarr` remain supported.
47
- - Avoid deprecated APIs: `ad.read`, `AnnData.concatenate()`, `AnnData.*_keys()`, and `anndata.__version__`. Prefer `ad.read_h5ad`, `ad.concat`, mapping `.keys()`, and `importlib.metadata.version("anndata")`.
48
- - Treat `anndata.experimental` APIs as useful but unstable. Prefer them for large-data workflows only when their current caveats are acceptable.
49
-
50
- ## Quick Start
51
-
52
- ### Creating an AnnData object
53
- ```python
54
- import anndata as ad
55
- import numpy as np
56
- import pandas as pd
57
-
58
- # Minimal creation
59
- X = np.random.rand(100, 2000) # 100 cells × 2000 genes
60
- adata = ad.AnnData(X)
61
-
62
- # With metadata
63
- obs = pd.DataFrame({
64
- 'cell_type': ['T cell', 'B cell'] * 50,
65
- 'sample': ['A', 'B'] * 50
66
- }, index=[f'cell_{i}' for i in range(100)])
67
-
68
- var = pd.DataFrame({
69
- 'gene_name': [f'Gene_{i}' for i in range(2000)]
70
- }, index=[f'ENSG{i:05d}' for i in range(2000)])
71
-
72
- adata = ad.AnnData(X=X, obs=obs, var=var)
73
- ```
74
-
75
- ### Reading data
76
- ```python
77
- # Native formats (read_h5ad/read_zarr remain at top-level)
78
- adata = ad.read_h5ad('data.h5ad')
79
- adata = ad.read_h5ad('large_data.h5ad', backed='r') # lazy load for large files
80
- adata = ad.read_zarr('data.zarr')
81
-
82
- # Other formats: prefer anndata.io (top-level imports are deprecated)
83
- from anndata.io import read_csv, read_loom, read_mtx
84
-
85
- adata = read_csv('data.csv')
86
- adata = read_loom('data.loom')
87
-
88
- # 10X Genomics: use scanpy (not anndata) — see scanpy skill
89
- import scanpy as sc
90
- adata = sc.read_10x_h5('filtered_feature_bc_matrix.h5')
91
- adata = sc.read_10x_mtx('filtered_feature_bc_matrix/')
92
- ```
93
-
94
- ### Writing data
95
- ```python
96
- # Write h5ad file
97
- adata.write_h5ad('output.h5ad')
98
-
99
- # Write with compression
100
- adata.write_h5ad('output.h5ad', compression='gzip')
101
-
102
- # Write other formats
103
- adata.write_zarr('output.zarr')
104
- adata.write_csvs('output_dir/')
105
- ```
106
-
107
- ### Basic operations
108
- ```python
109
- # Subset by conditions
110
- t_cells = adata[adata.obs['cell_type'] == 'T cell']
111
-
112
- # Subset by indices
113
- subset = adata[0:50, 0:100]
114
-
115
- # Add metadata
116
- adata.obs['quality_score'] = np.random.rand(adata.n_obs)
117
- adata.var['highly_variable'] = np.random.rand(adata.n_vars) > 0.8
118
-
119
- # Access dimensions
120
- print(f"{adata.n_obs} observations × {adata.n_vars} variables")
121
- ```
122
-
123
- ## Core Capabilities
124
-
125
- ### 1. Data Structure
126
-
127
- Understand the AnnData object structure including X, obs, var, layers, obsm, varm, obsp, varp, uns, and raw components.
128
-
129
- **See**: `references/data_structure.md` for comprehensive information on:
130
- - Core components (X, obs, var, layers, obsm, varm, obsp, varp, uns, raw)
131
- - Creating AnnData objects from various sources
132
- - Accessing and manipulating data components
133
- - Memory-efficient practices
134
-
135
- ### 2. Input/Output Operations
136
-
137
- Read and write data in various formats with support for compression, backed mode, and cloud storage.
138
-
139
- **See**: `references/io_operations.md` for details on:
140
- - Native formats (h5ad, zarr)
141
- - Alternative formats (CSV, MTX, Loom, 10X, Excel)
142
- - Backed mode for large datasets
143
- - Remote data access
144
- - Format conversion
145
- - Performance optimization
146
-
147
- Common commands:
148
- ```python
149
- from anndata.io import read_mtx
150
-
151
- # Read/write h5ad
152
- adata = ad.read_h5ad('data.h5ad', backed='r')
153
- adata.write_h5ad('output.h5ad', compression='gzip')
154
-
155
- # 10X Genomics (via scanpy)
156
- import scanpy as sc
157
- adata = sc.read_10x_h5('filtered_feature_bc_matrix.h5')
158
-
159
- # Read MTX format
160
- adata = read_mtx('matrix.mtx').T
161
- ```
162
-
163
- ### 3. Concatenation
164
-
165
- Combine multiple AnnData objects along observations or variables with flexible join strategies.
166
-
167
- **See**: `references/concatenation.md` for comprehensive coverage of:
168
- - Basic concatenation (axis=0 for observations, axis=1 for variables)
169
- - Join types (inner, outer)
170
- - Merge strategies (same, unique, first, only)
171
- - Tracking data sources with labels
172
- - Lazy concatenation (AnnCollection)
173
- - On-disk concatenation for large datasets
174
-
175
- Common commands:
176
- ```python
177
- # Concatenate observations (combine samples)
178
- adata = ad.concat(
179
- [adata1, adata2, adata3],
180
- axis=0,
181
- join='inner',
182
- label='batch',
183
- keys=['batch1', 'batch2', 'batch3']
184
- )
185
-
186
- # Concatenate variables (combine modalities)
187
- adata = ad.concat([adata_rna, adata_protein], axis=1)
188
-
189
- # Lazy collection over backed AnnData objects (experimental)
190
- from anndata.experimental import AnnCollection
191
-
192
- backed_adatas = [
193
- ad.read_h5ad(path, backed='r')
194
- for path in ['data1.h5ad', 'data2.h5ad']
195
- ]
196
- collection = AnnCollection(
197
- backed_adatas,
198
- join_obs='outer',
199
- join_vars='inner',
200
- label='dataset'
201
- )
202
- ```
203
-
204
- ### 4. Data Manipulation
205
-
206
- Transform, subset, filter, and reorganize data efficiently.
207
-
208
- **See**: `references/manipulation.md` for detailed guidance on:
209
- - Subsetting (by indices, names, boolean masks, metadata conditions)
210
- - Transposition
211
- - Copying (full copies vs views)
212
- - Renaming (observations, variables, categories)
213
- - Type conversions (strings to categoricals, sparse/dense)
214
- - Adding/removing data components
215
- - Reordering
216
- - Quality control filtering
217
-
218
- Common commands:
219
- ```python
220
- # Subset by metadata
221
- filtered = adata[adata.obs['quality_score'] > 0.8]
222
- hv_genes = adata[:, adata.var['highly_variable']]
223
-
224
- # Transpose
225
- adata_T = adata.T
226
-
227
- # Copy vs view
228
- view = adata[0:100, :] # View (lightweight reference)
229
- copy = adata[0:100, :].copy() # Independent copy
230
-
231
- # Convert strings to categoricals
232
- adata.strings_to_categoricals()
233
- ```
234
-
235
- ### 5. Best Practices
236
-
237
- Follow recommended patterns for memory efficiency, performance, and reproducibility.
238
-
239
- **See**: `references/best_practices.md` for guidelines on:
240
- - Memory management (sparse matrices, categoricals, backed mode)
241
- - Views vs copies
242
- - Data storage optimization
243
- - Performance optimization
244
- - Working with raw data
245
- - Metadata management
246
- - Reproducibility
247
- - Error handling
248
- - Integration with other tools
249
- - Common pitfalls and solutions
250
-
251
- Key recommendations:
252
- ```python
253
- # Use sparse matrices for sparse data
254
- from scipy.sparse import csr_matrix
255
- adata.X = csr_matrix(adata.X)
256
-
257
- # Convert strings to categoricals
258
- adata.strings_to_categoricals()
259
-
260
- # Use backed mode for large files
261
- adata = ad.read_h5ad('large.h5ad', backed='r')
262
-
263
- # Store raw before filtering
264
- adata.raw = adata.copy()
265
- adata = adata[:, adata.var['highly_variable']]
266
- ```
267
-
268
- ## Integration with Scverse Ecosystem
269
-
270
- AnnData serves as the foundational data structure for the scverse ecosystem:
271
-
272
- ### Scanpy (Single-cell analysis)
273
- ```python
274
- import scanpy as sc
275
-
276
- # Preprocessing
277
- sc.pp.filter_cells(adata, min_genes=200)
278
- sc.pp.normalize_total(adata, target_sum=1e4)
279
- sc.pp.log1p(adata)
280
- sc.pp.highly_variable_genes(adata, n_top_genes=2000)
281
-
282
- # Dimensionality reduction
283
- sc.pp.pca(adata, n_comps=50)
284
- sc.pp.neighbors(adata, n_neighbors=15)
285
- sc.tl.umap(adata)
286
- sc.tl.leiden(adata)
287
-
288
- # Visualization
289
- sc.pl.umap(adata, color=['cell_type', 'leiden'])
290
- ```
291
-
292
- ### Muon (Multimodal data)
293
- ```python
294
- import muon as mu
295
-
296
- # Combine RNA and protein data
297
- mdata = mu.MuData({'rna': adata_rna, 'protein': adata_protein})
298
- ```
299
-
300
- ### PyTorch integration
301
- ```python
302
- from anndata.experimental import AnnLoader
303
-
304
- # Create DataLoader for deep learning
305
- dataloader = AnnLoader(adata, batch_size=128, shuffle=True)
306
-
307
- for batch in dataloader:
308
- X = batch.X
309
- # Train model
310
- ```
311
-
312
- ## Common Workflows
313
-
314
- ### Single-cell RNA-seq analysis
315
- ```python
316
- import anndata as ad
317
- import scanpy as sc
318
-
319
- # 1. Load data (10X via scanpy; anndata handles h5ad/zarr natively)
320
- adata = sc.read_10x_h5('filtered_feature_bc_matrix.h5')
321
-
322
- # 2. Quality control
323
- adata.obs['n_genes'] = (adata.X > 0).sum(axis=1)
324
- adata.obs['n_counts'] = adata.X.sum(axis=1)
325
- adata = adata[adata.obs['n_genes'] > 200]
326
- adata = adata[adata.obs['n_counts'] < 50000]
327
-
328
- # 3. Store raw
329
- adata.raw = adata.copy()
330
-
331
- # 4. Normalize and filter
332
- sc.pp.normalize_total(adata, target_sum=1e4)
333
- sc.pp.log1p(adata)
334
- sc.pp.highly_variable_genes(adata, n_top_genes=2000)
335
- adata = adata[:, adata.var['highly_variable']]
336
-
337
- # 5. Save processed data
338
- adata.write_h5ad('processed.h5ad')
339
- ```
340
-
341
- ### Batch integration
342
- ```python
343
- # Load multiple batches
344
- adata1 = ad.read_h5ad('batch1.h5ad')
345
- adata2 = ad.read_h5ad('batch2.h5ad')
346
- adata3 = ad.read_h5ad('batch3.h5ad')
347
-
348
- # Concatenate with batch labels
349
- adata = ad.concat(
350
- [adata1, adata2, adata3],
351
- label='batch',
352
- keys=['batch1', 'batch2', 'batch3'],
353
- join='inner'
354
- )
355
-
356
- # Apply batch correction
357
- import scanpy as sc
358
- sc.pp.combat(adata, key='batch')
359
-
360
- # Continue analysis
361
- sc.pp.pca(adata)
362
- sc.pp.neighbors(adata)
363
- sc.tl.umap(adata)
364
- ```
365
-
366
- ### Working with large datasets
367
- ```python
368
- # Open in backed mode
369
- adata = ad.read_h5ad('100GB_dataset.h5ad', backed='r')
370
-
371
- # Filter based on metadata (no data loading)
372
- high_quality = adata[adata.obs['quality_score'] > 0.8]
373
-
374
- # Load filtered subset
375
- adata_subset = high_quality.to_memory()
376
-
377
- # Process subset
378
- process(adata_subset)
379
-
380
- # Or process in chunks
381
- chunk_size = 1000
382
- for i in range(0, adata.n_obs, chunk_size):
383
- chunk = adata[i:i+chunk_size, :].to_memory()
384
- process(chunk)
385
- ```
386
-
387
- ## Troubleshooting
388
-
389
- ### Out of memory errors
390
- Use backed mode or convert to sparse matrices:
391
- ```python
392
- # Backed mode
393
- adata = ad.read_h5ad('file.h5ad', backed='r')
394
-
395
- # Sparse matrices
396
- from scipy.sparse import csr_matrix
397
- adata.X = csr_matrix(adata.X)
398
- ```
399
-
400
- ### Slow file reading
401
- Use compression and appropriate formats:
402
- ```python
403
- # Optimize for storage
404
- adata.strings_to_categoricals()
405
- adata.write_h5ad('file.h5ad', compression='gzip')
406
-
407
- # Use Zarr for cloud storage; v3 writes are opt-in in anndata 0.12
408
- import anndata as ad
409
-
410
- ad.settings.zarr_write_format = 3
411
- ad.settings.auto_shard_zarr_v3 = True # experimental; independent of zarr_write_format
412
- adata.write_zarr('file.zarr', chunks=(1000, 1000))
413
- ```
414
-
415
- ### Index alignment issues
416
- Always align external data on index:
417
- ```python
418
- # Wrong
419
- adata.obs['new_col'] = external_data['values']
420
-
421
- # Correct
422
- adata.obs['new_col'] = external_data.set_index('cell_id').loc[adata.obs_names, 'values']
423
- ```
424
-
425
- ## Additional Resources
426
-
427
- - **Official documentation**: https://anndata.readthedocs.io/
428
- - **Scanpy tutorials**: https://scanpy.readthedocs.io/
429
- - **Scverse ecosystem**: https://scverse.org/
430
- - **GitHub repository**: https://github.com/scverse/anndata
431
-
@@ -1,152 +0,0 @@
1
- ---
2
- name: arbor
3
- description: Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
4
- allowed-tools: Read Write Edit Bash Agent
5
- license: MIT license
6
- metadata:
7
- version: "1.1"
8
- skill-author: K-Dense Inc.
9
- ---
10
-
11
- # Arbor — Autonomous Optimization via Hypothesis Tree Refinement
12
-
13
- ## Overview
14
-
15
- This skill runs an **Autonomous Optimization (AO)** loop: starting from an existing artifact and a measurable objective, improve it through many rounds of experiment and evaluation — without step-by-step human supervision and without overfitting to the feedback signal. It's the right tool when the bottleneck isn't writing one good change, but *organizing dozens of trials* so that lessons accumulate instead of evaporating.
16
-
17
- It implements **Hypothesis Tree Refinement (HTR)** from *Arbor* (Jin et al., 2026). The key idea: keep the research state in a persistent **hypothesis tree** rather than in conversation history. Each node binds a hypothesis, the distilled insight it produced, and a pointer to the artifact version that realizes it. You play the long-lived **coordinator** that owns this tree and decides where to search; short-lived **executor** subagents test one hypothesis each in isolated git worktrees and report back. A **held-out merge gate** admits a change only when it improves on a *test* evaluator the search never optimized against. This is what turns trial-and-error into cumulative, auditable research.
18
-
19
- Use the `scripts/tree.py` state manager for all the bookkeeping (creating nodes, writing evidence, propagating insights, pruning, the merge gate, the Observe projection). It keeps the state consistent and frees you to spend judgment on what the evidence *means*.
20
-
21
- ## When to use this skill
22
-
23
- Reach for Arbor when the task is **iterative improvement of a concrete artifact under an evaluator**:
24
- - Model training: optimizer/architecture/recipe changes to lower loss or hit a target in fewer steps.
25
- - Harness/agent engineering: raising pass rate or accuracy of an agent loop, search harness, or tool-use scaffold.
26
- - Data synthesis: improving a generation/filtering pipeline judged by downstream model behavior.
27
- - Benchmark optimization: MLE-bench / Kaggle-style "improve the submission" tasks.
28
- - Prompt/system optimization where you can score outputs automatically.
29
-
30
- The distinguishing signals: there's an **artifact you can modify**, an **objective**, a way to **score** candidates, and you expect to run **many experiments**. If the user only wants a single fix or a one-shot answer, this is overkill — just do the work directly. If they want open-ended ideation with no evaluator, use `hypothesis-generation` or `scientific-brainstorming` instead.
31
-
32
- ## The AO setup — pin this down first
33
-
34
- Before any experiments, establish the task tuple `(M_0, O, E_dev, E_test)`. Getting this right matters more than any later decision, so confirm it explicitly:
35
-
36
- - **M_0 — initial material**: the artifact to improve (a repo, a script, a config, a prompt). Make sure it's under git and currently runs.
37
- - **O — objective**: the natural-language goal and the metric *direction* (maximize accuracy? minimize loss/steps?).
38
- - **E_dev — development evaluator**: a command you can run freely during search to score a candidate. Fast, repeatable.
39
- - **E_test — held-out test evaluator**: a *separate* evaluator (different seeds, different split, or a larger run) used only at the merge gate. It must not be used as a search oracle — that's the whole point.
40
-
41
- If the user hasn't given you a clean dev/test split, **construct one and say so**. The dev/test separation is the mechanism that catches overfitting: a candidate that wins on dev but not on test isn't a success, it's a warning that you're exploiting the feedback signal. Without it, autonomous search reliably overfits.
42
-
43
- Initialize the run:
44
-
45
- ```bash
46
- python scripts/tree.py init \
47
- --objective "Improve BrowseComp answer accuracy on the search harness" \
48
- --dev-eval "python eval.py --split dev --n 50" \
49
- --test-eval "python eval.py --split test --n 300" \
50
- --material "." --metric-direction max --branching 3 --max-depth 2 --budget 12
51
- ```
52
-
53
- `--branching` is how many sibling hypotheses you propose per parent; `--max-depth 2` keeps directions at depth 1 and concrete interventions at depth 2 (the paper's default); `--budget` is the number of coordinator cycles. Start small (10–20 cycles) — structured search beats brute force, and you can extend if progress is still being made.
54
-
55
- ## The coordinator loop
56
-
57
- You run repeated cycles of six steps. This is the heart of HTR; do not collapse it into ad-hoc editing. Run `python scripts/tree.py cycle` once per cycle to track the budget.
58
-
59
- ### 1. Observe
60
- Begin every cycle by re-grounding in the tree, not in your memory of the conversation:
61
-
62
- ```bash
63
- python scripts/tree.py observe
64
- ```
65
-
66
- This prints the objective, global insights, the active frontier (selectable hypotheses), executed nodes with their evidence, pruned lessons (negative constraints), and the current best artifact. Treating the tree as the source of truth is what keeps you coherent over a long run, after context compression has thrown away the details.
67
-
68
- ### 2. Ideate
69
- Pick a promising parent and propose a few child hypotheses under it. **Condition on the tree's evidence** — this is the difference between Arbor and random search:
70
- - Validated insights are assumptions you can build on.
71
- - Pruned nodes are dead ends to avoid.
72
- - A "half-right" result is a *starting point for a sharper hypothesis*, not a reason to abandon the direction.
73
-
74
- Each hypothesis should be a **falsifiable claim about how changing the artifact will move the metric**, not a vague intention. Depth-1 nodes are broad directions ("the search harness loses correct answers it already retrieved"); depth-2 nodes are concrete, executable interventions ("run K=5 independent rollouts and aggregate by evidence dossier instead of majority vote").
75
-
76
- ```bash
77
- python scripts/tree.py add-node --parent n0 --hypothesis "Verification, not retrieval, is the bottleneck: candidates are found but discarded"
78
- python scripts/tree.py add-node --parent n4 --hypothesis "Decompose the question into atomic constraints and verify each independently"
79
- ```
80
-
81
- ### 3. Select
82
- Choose which pending leaves to run next. **Selection is not pure score-maximization** — pick a hypothesis because it has strong prior evidence, because it would resolve an ambiguity its siblings exposed, or because its failure would clarify an important assumption. Frontier control under delayed feedback rewards informative experiments, not just promising ones.
83
-
84
- ### 4. Dispatch
85
- Run each selected hypothesis as an **executor subagent in an isolated worktree** (use the Agent tool with `isolation: "worktree"`, or have the executor create one with `git worktree add`). Isolation matters: parallel experiments must not clobber each other or the current best, and exploratory changes stay quarantined until they pass the merge gate.
86
-
87
- Dispatch siblings **in parallel** (multiple Agent calls in one message) when they're independent — comparative evidence within one direction is exactly what makes later pruning and abstraction possible.
88
-
89
- Give each executor a tight, **hypothesis-bound** brief. See `references/executor-brief.md` for the full template. The contract that makes HTR work: **the executor may not change the hypothesis when the metric stalls.** It repairs its own code and reruns, but `h_n` is fixed — otherwise the returned score is no longer evidence about the assigned node and the tree's semantics break. The executor returns exactly four things:
90
- - **dev_score** — the dev evaluator result (for selection);
91
- - **result** — a factual summary of what happened;
92
- - **insight** — the distilled, reusable lesson (*why* the result supports, weakens, or bounds the hypothesis);
93
- - **branch_ref** — the git branch/commit/worktree path holding the artifact.
94
-
95
- Mark a node `running` before dispatch (`tree.py set-status --node n5 --status running`) so the Observe projection stays accurate.
96
-
97
- ### 5. Backpropagate
98
- When an executor returns, write its report into the node, then **abstract the lesson upward**:
99
-
100
- ```bash
101
- python scripts/tree.py set-evidence --node n5 --dev-score 70.0 \
102
- --result "K=5 dossier aggregation recovers answers in minority rollouts" \
103
- --insight "Correct answers often appear in a minority of rollouts; aggregation beats majority vote" \
104
- --branch-ref "wt/n5"
105
-
106
- python scripts/tree.py propagate --node n5 \
107
- --insight "Candidate coverage, not verification, limits this direction" --to-root
108
- ```
109
-
110
- This is the step that makes the tree more than a log. A leaf-level observation ("data-interface mismatch") should become a direction-level constraint and, if it generalizes, a global prior that shapes future ideation. **Insight propagation is the component that drives most of HTR's gains** — in the paper's MLE-Bench Lite ablation, a tree *without* insight feedback scored even lower than a flat experiment queue with no tree at all (54.5% vs. 63.6% any-medal, against 81.8% for the full system). Hierarchy alone isn't enough: the semantic memory is what matters. So spend real thought on the abstraction; don't just copy the leaf insight upward verbatim.
111
-
112
- ### 6. Decide
113
- Decide what to do with the new evidence: keep expanding a direction, prune a falsified subtree, or attempt to merge a candidate.
114
-
115
- - **Prune** dead ends, recording *why* — the reason becomes a negative constraint:
116
- ```bash
117
- python scripts/tree.py prune --node n7 --reason "search-augmented judge overfits dev questions; no test transfer"
118
- ```
119
- - **Merge gate** — promote a candidate to the new best **only if it improves on `E_test`**. Run the test evaluator in a *fresh* worktree (not the dev worktree, to avoid leakage), then:
120
- ```bash
121
- python scripts/tree.py merge --node n5 --test-score 67.67 --branch-ref "wt/n5"
122
- ```
123
- If the gate rejects it, that's informative: a high-dev / low-test candidate is evidence the direction may be exploiting the dev signal rather than producing a transferable improvement. Record that lesson; don't quietly promote it anyway.
124
-
125
- Repeat until the budget is spent, the frontier is exhausted, or progress has clearly stalled.
126
-
127
- ## Finishing the run
128
-
129
- When you stop, produce a short report (see `references/report-template.md`) covering:
130
- - the final best artifact, its test score, and its delta over `M_0`;
131
- - the tree (`python scripts/tree.py status`) as the audit trail of what was tried;
132
- - the main hypothesis shifts — how task understanding deepened across the run (early nodes test broad mechanisms; later nodes find their limits; ancestor insights compress these into the constraints behind the final design);
133
- - merged vs. explored: many nodes improve dev, far fewer pass the test gate — report that gap honestly rather than overstating dev wins.
134
-
135
- Always leave `M_best` as a real, runnable artifact on a named branch, and tell the user how to check it out.
136
-
137
- ## Principles that make this work (not rote rules)
138
-
139
- These come from the paper's analysis; understanding *why* matters more than following them mechanically.
140
-
141
- - **The tree is the memory; conversation is not.** Over a long horizon your context gets compressed. Re-Observe each cycle so decisions rest on durable evidence, not a lossy summary.
142
- - **Structured search, not more sampling.** Arbor's gains come from how the budget is *organized* — maintaining competing hypotheses, comparing siblings, carrying lessons forward — not from spending more tokens. Don't fan out aimlessly; each experiment should be conditioned on what the tree already knows.
143
- - **Dev guides, test admits.** Use dev feedback freely to steer exploration, but never let a dev win into the final artifact without test confirmation. The dev/test disagreement is itself a signal worth reading.
144
- - **Executors are hypothesis-bound.** Local engineering flexibility (edit, debug, rerun) is fine; silently changing the hypothesis to chase a better number is not — it destroys the meaning of the evidence.
145
- - **Failures are constraints, not noise.** A falsified hypothesis tells you what the solution must avoid. Pruned-with-a-reason is more valuable than pruned-and-forgotten.
146
-
147
- ## Reference files
148
-
149
- - `references/htr-methodology.md` — deeper explanation of HTR, the node structure, the six steps, and the paper's empirical lessons (ablations, transfer, cost). Read when you want the rationale behind a design choice.
150
- - `references/executor-brief.md` — the template for the brief you hand each executor subagent.
151
- - `references/report-template.md` — the final-report structure.
152
- - `references/arbor-upstream.md` — how to install and run the standalone `arbor` CLI from RUC-NLPIR/Arbor instead of orchestrating it natively, and when to prefer each.