@pikaa-ai/pikaa 0.3.23 → 0.3.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +407 -219
- package/dist/index.js +7 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
package/skills/arboreto/SKILL.md
DELETED
|
@@ -1,267 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: arboreto
|
|
3
|
-
description: Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
|
|
4
|
-
license: BSD-3-Clause license
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.0"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Arboreto
|
|
11
|
-
|
|
12
|
-
## Overview
|
|
13
|
-
|
|
14
|
-
Arboreto is a Python library from [Aerts Lab](https://github.com/aertslab/arboreto) for inferring gene regulatory networks (GRNs) from gene expression data. It parallelizes tree-based ensemble regression (GRNBoost2, GENIE3) with [Dask](https://distributed.dask.org/) across local cores or remote clusters.
|
|
15
|
-
|
|
16
|
-
**Core capability**: Identify which transcription factors (TFs) regulate which target genes based on expression patterns across observations (cells, samples, conditions).
|
|
17
|
-
|
|
18
|
-
**Upstream**: PyPI **0.1.6** (2021-02-09, latest). Docs: [arboreto.readthedocs.io](https://arboreto.readthedocs.io/en/latest/). Primary downstream consumer: [pySCENIC](https://github.com/aertslab/pySCENIC).
|
|
19
|
-
|
|
20
|
-
## Quick Start
|
|
21
|
-
|
|
22
|
-
Install arboreto:
|
|
23
|
-
```bash
|
|
24
|
-
uv pip install arboreto
|
|
25
|
-
```
|
|
26
|
-
|
|
27
|
-
Basic GRN inference:
|
|
28
|
-
```python
|
|
29
|
-
import pandas as pd
|
|
30
|
-
from arboreto.algo import grnboost2
|
|
31
|
-
|
|
32
|
-
if __name__ == '__main__':
|
|
33
|
-
# Load expression data (genes as columns)
|
|
34
|
-
expression_matrix = pd.read_csv('expression_data.tsv', sep='\t')
|
|
35
|
-
|
|
36
|
-
# Infer regulatory network
|
|
37
|
-
network = grnboost2(expression_data=expression_matrix)
|
|
38
|
-
|
|
39
|
-
# Save results (TF, target, importance)
|
|
40
|
-
network.to_csv('network.tsv', sep='\t', index=False, header=False)
|
|
41
|
-
```
|
|
42
|
-
|
|
43
|
-
**Critical**: Always use `if __name__ == '__main__':` guard because Dask spawns new processes.
|
|
44
|
-
|
|
45
|
-
## Core Capabilities
|
|
46
|
-
|
|
47
|
-
### 1. Basic GRN Inference
|
|
48
|
-
|
|
49
|
-
For standard GRN inference workflows including:
|
|
50
|
-
- Input data preparation (Pandas DataFrame or NumPy array)
|
|
51
|
-
- Running inference with GRNBoost2 or GENIE3
|
|
52
|
-
- Filtering by transcription factors
|
|
53
|
-
- Output format and interpretation
|
|
54
|
-
|
|
55
|
-
**See**: `references/basic_inference.md`
|
|
56
|
-
|
|
57
|
-
**Use the ready-to-run script**: `scripts/basic_grn_inference.py` for standard inference tasks:
|
|
58
|
-
```bash
|
|
59
|
-
python scripts/basic_grn_inference.py expression_data.tsv output_network.tsv --tf-file tfs.txt --seed 777 --limit 5000
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
### 2. Algorithm Selection
|
|
63
|
-
|
|
64
|
-
Arboreto provides two algorithms:
|
|
65
|
-
|
|
66
|
-
**GRNBoost2 (Recommended)**:
|
|
67
|
-
- Fast gradient boosting-based inference
|
|
68
|
-
- Optimized for large datasets (10k+ observations)
|
|
69
|
-
- Default choice for most analyses
|
|
70
|
-
|
|
71
|
-
**GENIE3**:
|
|
72
|
-
- Random Forest-based inference
|
|
73
|
-
- Original multiple regression approach
|
|
74
|
-
- Use for comparison or validation
|
|
75
|
-
|
|
76
|
-
Quick comparison:
|
|
77
|
-
```python
|
|
78
|
-
from arboreto.algo import grnboost2, genie3
|
|
79
|
-
|
|
80
|
-
# Fast, recommended
|
|
81
|
-
network_grnboost = grnboost2(expression_data=matrix)
|
|
82
|
-
|
|
83
|
-
# Classic algorithm
|
|
84
|
-
network_genie3 = genie3(expression_data=matrix)
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
**For detailed algorithm comparison, parameters, and selection guidance**: `references/algorithms.md`
|
|
88
|
-
|
|
89
|
-
### 3. Distributed Computing
|
|
90
|
-
|
|
91
|
-
Scale inference from local multi-core to cluster environments:
|
|
92
|
-
|
|
93
|
-
**Local (default)** - Uses all available cores automatically:
|
|
94
|
-
```python
|
|
95
|
-
network = grnboost2(expression_data=matrix)
|
|
96
|
-
```
|
|
97
|
-
|
|
98
|
-
**Custom local client** - Control resources:
|
|
99
|
-
```python
|
|
100
|
-
from distributed import LocalCluster, Client
|
|
101
|
-
|
|
102
|
-
local_cluster = LocalCluster(n_workers=10, memory_limit='8GB')
|
|
103
|
-
client = Client(local_cluster)
|
|
104
|
-
|
|
105
|
-
network = grnboost2(expression_data=matrix, client_or_address=client)
|
|
106
|
-
|
|
107
|
-
client.close()
|
|
108
|
-
local_cluster.close()
|
|
109
|
-
```
|
|
110
|
-
|
|
111
|
-
**Cluster computing** - Connect to remote Dask scheduler:
|
|
112
|
-
```python
|
|
113
|
-
from distributed import Client
|
|
114
|
-
|
|
115
|
-
client = Client('tcp://scheduler:8786')
|
|
116
|
-
network = grnboost2(expression_data=matrix, client_or_address=client)
|
|
117
|
-
```
|
|
118
|
-
|
|
119
|
-
**For cluster setup, performance optimization, and large-scale workflows**: `references/distributed_computing.md`
|
|
120
|
-
|
|
121
|
-
## Installation
|
|
122
|
-
|
|
123
|
-
```bash
|
|
124
|
-
uv pip install arboreto
|
|
125
|
-
```
|
|
126
|
-
|
|
127
|
-
Conda (Bioconda):
|
|
128
|
-
|
|
129
|
-
```bash
|
|
130
|
-
conda install -c bioconda arboreto
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
**Dependencies** (from upstream `requirements.txt`): `dask[complete]`, `distributed`, `numpy`, `pandas`, `scikit-learn`, `scipy`
|
|
134
|
-
|
|
135
|
-
**Input formats**: pandas DataFrame, dense `numpy.ndarray`, or sparse `scipy.sparse.csc_matrix` (rows = observations, columns = genes). For array/matrix inputs, pass `gene_names` explicitly.
|
|
136
|
-
|
|
137
|
-
## Common Use Cases
|
|
138
|
-
|
|
139
|
-
### Single-Cell RNA-seq Analysis
|
|
140
|
-
```python
|
|
141
|
-
import pandas as pd
|
|
142
|
-
from arboreto.algo import grnboost2
|
|
143
|
-
|
|
144
|
-
if __name__ == '__main__':
|
|
145
|
-
# Load single-cell expression matrix (cells x genes)
|
|
146
|
-
sc_data = pd.read_csv('scrna_counts.tsv', sep='\t')
|
|
147
|
-
|
|
148
|
-
# Infer cell-type-specific regulatory network
|
|
149
|
-
network = grnboost2(expression_data=sc_data, seed=42)
|
|
150
|
-
|
|
151
|
-
# Filter high-confidence links
|
|
152
|
-
high_confidence = network[network['importance'] > 0.5]
|
|
153
|
-
high_confidence.to_csv('grn_high_confidence.tsv', sep='\t', index=False)
|
|
154
|
-
```
|
|
155
|
-
|
|
156
|
-
### Bulk RNA-seq with TF Filtering
|
|
157
|
-
```python
|
|
158
|
-
from arboreto.utils import load_tf_names
|
|
159
|
-
from arboreto.algo import grnboost2
|
|
160
|
-
|
|
161
|
-
if __name__ == '__main__':
|
|
162
|
-
# Load data
|
|
163
|
-
expression_data = pd.read_csv('rnaseq_tpm.tsv', sep='\t')
|
|
164
|
-
tf_names = load_tf_names('human_tfs.txt')
|
|
165
|
-
|
|
166
|
-
# Infer with TF restriction
|
|
167
|
-
network = grnboost2(
|
|
168
|
-
expression_data=expression_data,
|
|
169
|
-
tf_names=tf_names,
|
|
170
|
-
seed=123
|
|
171
|
-
)
|
|
172
|
-
|
|
173
|
-
network.to_csv('tf_target_network.tsv', sep='\t', index=False)
|
|
174
|
-
```
|
|
175
|
-
|
|
176
|
-
### Comparative Analysis (Multiple Conditions)
|
|
177
|
-
```python
|
|
178
|
-
from arboreto.algo import grnboost2
|
|
179
|
-
|
|
180
|
-
if __name__ == '__main__':
|
|
181
|
-
# Infer networks for different conditions
|
|
182
|
-
conditions = ['control', 'treatment_24h', 'treatment_48h']
|
|
183
|
-
|
|
184
|
-
for condition in conditions:
|
|
185
|
-
data = pd.read_csv(f'{condition}_expression.tsv', sep='\t')
|
|
186
|
-
network = grnboost2(expression_data=data, seed=42)
|
|
187
|
-
network.to_csv(f'{condition}_network.tsv', sep='\t', index=False)
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
## Output Interpretation
|
|
191
|
-
|
|
192
|
-
Arboreto returns a DataFrame with regulatory links:
|
|
193
|
-
|
|
194
|
-
| Column | Description |
|
|
195
|
-
|--------|-------------|
|
|
196
|
-
| `TF` | Transcription factor (regulator) |
|
|
197
|
-
| `target` | Target gene |
|
|
198
|
-
| `importance` | Regulatory importance score (higher = stronger) |
|
|
199
|
-
|
|
200
|
-
**Filtering strategy**:
|
|
201
|
-
- `limit=N` at inference time (return top N links globally)
|
|
202
|
-
- Post-hoc importance threshold (e.g., > 0.5)
|
|
203
|
-
- Top links per target via `groupby('target')`
|
|
204
|
-
- Statistical significance testing (permutation tests, external tools)
|
|
205
|
-
|
|
206
|
-
## Integration with pySCENIC
|
|
207
|
-
|
|
208
|
-
Arboreto powers the GRN inference step in [pySCENIC](https://github.com/aertslab/pySCENIC). pySCENIC 0.11+ passes sparse expression matrices to `grnboost2` / `genie3`; pySCENIC 0.12+ defaults to `arboreto_with_multiprocessing.py` (no Dask) for compatibility — use standalone arboreto when you need Dask scaling.
|
|
209
|
-
|
|
210
|
-
```python
|
|
211
|
-
# Standalone: infer co-expression modules before pySCENIC cisTarget pruning
|
|
212
|
-
from arboreto.algo import grnboost2
|
|
213
|
-
|
|
214
|
-
network = grnboost2(expression_data=expression_df, tf_names=tf_list, limit=5000)
|
|
215
|
-
|
|
216
|
-
# Downstream: pySCENIC ctx pruning, regulon definition, AUCell (see pySCENIC docs)
|
|
217
|
-
```
|
|
218
|
-
|
|
219
|
-
Convert AnnData to a DataFrame for arboreto directly:
|
|
220
|
-
|
|
221
|
-
```python
|
|
222
|
-
expression_df = adata.to_df() # cells x genes
|
|
223
|
-
```
|
|
224
|
-
|
|
225
|
-
## Reproducibility
|
|
226
|
-
|
|
227
|
-
Always set a seed for reproducible results:
|
|
228
|
-
```python
|
|
229
|
-
network = grnboost2(expression_data=matrix, seed=777)
|
|
230
|
-
```
|
|
231
|
-
|
|
232
|
-
Run multiple seeds for robustness analysis:
|
|
233
|
-
```python
|
|
234
|
-
from distributed import LocalCluster, Client
|
|
235
|
-
|
|
236
|
-
if __name__ == '__main__':
|
|
237
|
-
client = Client(LocalCluster())
|
|
238
|
-
|
|
239
|
-
seeds = [42, 123, 777]
|
|
240
|
-
networks = []
|
|
241
|
-
|
|
242
|
-
for seed in seeds:
|
|
243
|
-
net = grnboost2(expression_data=matrix, client_or_address=client, seed=seed)
|
|
244
|
-
networks.append(net)
|
|
245
|
-
|
|
246
|
-
# Consensus: links recurring across runs (example: mean importance per TF-target pair)
|
|
247
|
-
import pandas as pd
|
|
248
|
-
combined = pd.concat(networks)
|
|
249
|
-
consensus = (
|
|
250
|
-
combined.groupby(['TF', 'target'], as_index=False)['importance']
|
|
251
|
-
.mean()
|
|
252
|
-
.query('importance > 0.5')
|
|
253
|
-
)
|
|
254
|
-
```
|
|
255
|
-
|
|
256
|
-
## Troubleshooting
|
|
257
|
-
|
|
258
|
-
**Memory errors**: Reduce dataset size by filtering low-variance genes or use distributed computing
|
|
259
|
-
|
|
260
|
-
**Slow performance**: Use GRNBoost2 instead of GENIE3, enable distributed client, filter TF list
|
|
261
|
-
|
|
262
|
-
**Dask errors**: Ensure `if __name__ == '__main__':` guard is present in scripts (required on Windows/macOS with spawn-based multiprocessing)
|
|
263
|
-
|
|
264
|
-
**Empty results**: Check data format (genes as columns), verify TF names match column names in the expression matrix
|
|
265
|
-
|
|
266
|
-
**Sparse data**: Use `scipy.sparse.csc_matrix` and pass matching `gene_names`; supported since arboreto 0.1.6 / pySCENIC 0.11
|
|
267
|
-
|
package/skills/astropy/SKILL.md
DELETED
|
@@ -1,353 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: astropy
|
|
3
|
-
description: Core Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
|
|
4
|
-
license: BSD-3-Clause license
|
|
5
|
-
compatibility: Requires Python 3.11+ with astropy installed (uv for package installation). Some features (object name resolution, site lookups, remote FITS reads, IERS updates) need network access.
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.2"
|
|
8
|
-
skill-author: K-Dense Inc.
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# Astropy
|
|
12
|
-
|
|
13
|
-
## Overview
|
|
14
|
-
|
|
15
|
-
Astropy is the core Python package for astronomy, providing essential functionality for astronomical research and data analysis. Use astropy for coordinate transformations, unit and quantity calculations, FITS file operations, cosmological calculations, precise time handling, tabular data manipulation, and astronomical image processing.
|
|
16
|
-
|
|
17
|
-
## When to Use This Skill
|
|
18
|
-
|
|
19
|
-
Use astropy when tasks involve:
|
|
20
|
-
- Converting between celestial coordinate systems (ICRS, Galactic, FK5, AltAz, etc.)
|
|
21
|
-
- Working with physical units and quantities (converting Jy to mJy, parsecs to km, etc.)
|
|
22
|
-
- Reading, writing, or manipulating FITS files (images or tables)
|
|
23
|
-
- Cosmological calculations (luminosity distance, lookback time, Hubble parameter)
|
|
24
|
-
- Precise time handling with different time scales (UTC, TAI, TT, TDB) and formats (JD, MJD, ISO)
|
|
25
|
-
- Table operations (reading catalogs, cross-matching, filtering, joining)
|
|
26
|
-
- WCS transformations between pixel and world coordinates
|
|
27
|
-
- Astronomical constants and calculations
|
|
28
|
-
|
|
29
|
-
## Quick Start
|
|
30
|
-
|
|
31
|
-
```python
|
|
32
|
-
import astropy.units as u
|
|
33
|
-
from astropy.coordinates import SkyCoord
|
|
34
|
-
from astropy.time import Time
|
|
35
|
-
from astropy.io import fits
|
|
36
|
-
from astropy.table import Table
|
|
37
|
-
from astropy.cosmology import Planck18
|
|
38
|
-
|
|
39
|
-
# Units and quantities
|
|
40
|
-
distance = 100 * u.pc
|
|
41
|
-
distance_km = distance.to(u.km)
|
|
42
|
-
|
|
43
|
-
# Coordinates
|
|
44
|
-
coord = SkyCoord(ra=10.5*u.degree, dec=41.2*u.degree, frame='icrs')
|
|
45
|
-
coord_galactic = coord.galactic
|
|
46
|
-
|
|
47
|
-
# Time
|
|
48
|
-
t = Time('2023-01-15 12:30:00')
|
|
49
|
-
jd = t.jd # Julian Date
|
|
50
|
-
|
|
51
|
-
# FITS files
|
|
52
|
-
data = fits.getdata('image.fits')
|
|
53
|
-
header = fits.getheader('image.fits')
|
|
54
|
-
|
|
55
|
-
# Tables
|
|
56
|
-
table = Table.read('catalog.fits')
|
|
57
|
-
|
|
58
|
-
# Cosmology
|
|
59
|
-
d_L = Planck18.luminosity_distance(z=1.0)
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
## Core Capabilities
|
|
63
|
-
|
|
64
|
-
### 1. Units and Quantities (`astropy.units`)
|
|
65
|
-
|
|
66
|
-
Handle physical quantities with units, perform unit conversions, and ensure dimensional consistency in calculations.
|
|
67
|
-
|
|
68
|
-
**Key operations:**
|
|
69
|
-
- Create quantities by multiplying values with units
|
|
70
|
-
- Convert between units using `.to()` method
|
|
71
|
-
- Perform arithmetic with automatic unit handling
|
|
72
|
-
- Use equivalencies for domain-specific conversions (spectral, doppler, parallax)
|
|
73
|
-
- Work with logarithmic units (magnitudes, decibels)
|
|
74
|
-
|
|
75
|
-
**See:** `references/units.md` for comprehensive documentation, unit systems, equivalencies, performance optimization, and unit arithmetic.
|
|
76
|
-
|
|
77
|
-
### 2. Coordinate Systems (`astropy.coordinates`)
|
|
78
|
-
|
|
79
|
-
Represent celestial positions and transform between different coordinate frames.
|
|
80
|
-
|
|
81
|
-
**Key operations:**
|
|
82
|
-
- Create coordinates with `SkyCoord` in any frame (ICRS, Galactic, FK5, AltAz, etc.)
|
|
83
|
-
- Transform between coordinate systems
|
|
84
|
-
- Calculate angular separations and position angles
|
|
85
|
-
- Match coordinates to catalogs
|
|
86
|
-
- Include distance for 3D coordinate operations
|
|
87
|
-
- Handle proper motions and radial velocities
|
|
88
|
-
- Query named objects from online databases
|
|
89
|
-
|
|
90
|
-
**See:** `references/coordinates.md` for detailed coordinate frame descriptions, transformations, observer-dependent frames (AltAz), catalog matching, and performance tips.
|
|
91
|
-
|
|
92
|
-
### 3. Cosmological Calculations (`astropy.cosmology`)
|
|
93
|
-
|
|
94
|
-
Perform cosmological calculations using standard cosmological models.
|
|
95
|
-
|
|
96
|
-
**Key operations:**
|
|
97
|
-
- Use built-in cosmologies (Planck18, WMAP9, etc.)
|
|
98
|
-
- Create custom cosmological models
|
|
99
|
-
- Calculate distances (luminosity, comoving, angular diameter)
|
|
100
|
-
- Compute ages and lookback times
|
|
101
|
-
- Determine Hubble parameter at any redshift
|
|
102
|
-
- Calculate density parameters and volumes
|
|
103
|
-
- Perform inverse calculations (find z for given distance)
|
|
104
|
-
|
|
105
|
-
**See:** `references/cosmology.md` for available models, distance calculations, time calculations, density parameters, and neutrino effects.
|
|
106
|
-
|
|
107
|
-
### 4. FITS File Handling (`astropy.io.fits`)
|
|
108
|
-
|
|
109
|
-
Read, write, and manipulate FITS (Flexible Image Transport System) files.
|
|
110
|
-
|
|
111
|
-
**Key operations:**
|
|
112
|
-
- Open FITS files with context managers
|
|
113
|
-
- Access HDUs (Header Data Units) by index or name
|
|
114
|
-
- Read and modify headers (keywords, comments, history)
|
|
115
|
-
- Work with image data (NumPy arrays)
|
|
116
|
-
- Handle table data (binary and ASCII tables)
|
|
117
|
-
- Create new FITS files (single or multi-extension)
|
|
118
|
-
- Use memory mapping for large files
|
|
119
|
-
- Access remote FITS files (S3, HTTP)
|
|
120
|
-
|
|
121
|
-
**See:** `references/fits.md` for comprehensive file operations, header manipulation, image and table handling, multi-extension files, and performance considerations.
|
|
122
|
-
|
|
123
|
-
### 5. Table Operations (`astropy.table`)
|
|
124
|
-
|
|
125
|
-
Work with tabular data with support for units, metadata, and various file formats.
|
|
126
|
-
|
|
127
|
-
**Key operations:**
|
|
128
|
-
- Create tables from arrays, lists, or dictionaries
|
|
129
|
-
- Read/write tables in multiple formats (FITS, CSV, HDF5, VOTable)
|
|
130
|
-
- Access and modify columns and rows
|
|
131
|
-
- Sort, filter, and index tables
|
|
132
|
-
- Perform database-style operations (join, group, aggregate)
|
|
133
|
-
- Stack and concatenate tables
|
|
134
|
-
- Work with unit-aware columns (QTable)
|
|
135
|
-
- Handle missing data with masking
|
|
136
|
-
|
|
137
|
-
**See:** `references/tables.md` for table creation, I/O operations, data manipulation, sorting, filtering, joins, grouping, and performance tips.
|
|
138
|
-
|
|
139
|
-
### 6. Time Handling (`astropy.time`)
|
|
140
|
-
|
|
141
|
-
Precise time representation and conversion between time scales and formats.
|
|
142
|
-
|
|
143
|
-
**Key operations:**
|
|
144
|
-
- Create Time objects in various formats (ISO, JD, MJD, Unix, etc.)
|
|
145
|
-
- Convert between time scales (UTC, TAI, TT, TDB, etc.)
|
|
146
|
-
- Perform time arithmetic with TimeDelta
|
|
147
|
-
- Calculate sidereal time for observers
|
|
148
|
-
- Compute light travel time corrections (barycentric, heliocentric)
|
|
149
|
-
- Work with time arrays efficiently
|
|
150
|
-
- Handle masked (missing) times
|
|
151
|
-
|
|
152
|
-
**See:** `references/time.md` for time formats, time scales, conversions, arithmetic, observing features, and precision handling.
|
|
153
|
-
|
|
154
|
-
### 7. World Coordinate System (`astropy.wcs`)
|
|
155
|
-
|
|
156
|
-
Transform between pixel coordinates in images and world coordinates.
|
|
157
|
-
|
|
158
|
-
**Key operations:**
|
|
159
|
-
- Read WCS from FITS headers
|
|
160
|
-
- Convert pixel coordinates to world coordinates (and vice versa)
|
|
161
|
-
- Calculate image footprints
|
|
162
|
-
- Access WCS parameters (reference pixel, projection, scale)
|
|
163
|
-
- Create custom WCS objects
|
|
164
|
-
|
|
165
|
-
**See:** `references/wcs_and_other_modules.md` for WCS operations and transformations.
|
|
166
|
-
|
|
167
|
-
## Additional Capabilities
|
|
168
|
-
|
|
169
|
-
The `references/wcs_and_other_modules.md` file also covers:
|
|
170
|
-
|
|
171
|
-
### NDData and CCDData
|
|
172
|
-
Containers for n-dimensional datasets with metadata, uncertainty, masking, and WCS information.
|
|
173
|
-
|
|
174
|
-
### Modeling
|
|
175
|
-
Framework for creating and fitting mathematical models to astronomical data.
|
|
176
|
-
|
|
177
|
-
### Visualization
|
|
178
|
-
Tools for astronomical image display with appropriate stretching and scaling.
|
|
179
|
-
|
|
180
|
-
### Constants
|
|
181
|
-
Physical and astronomical constants with proper units (speed of light, solar mass, Planck constant, etc.).
|
|
182
|
-
|
|
183
|
-
### Convolution
|
|
184
|
-
Image processing kernels for smoothing and filtering.
|
|
185
|
-
|
|
186
|
-
### Statistics
|
|
187
|
-
Robust statistical functions including sigma clipping and outlier rejection.
|
|
188
|
-
|
|
189
|
-
## Installation
|
|
190
|
-
|
|
191
|
-
```bash
|
|
192
|
-
# Reproducible install against the current stable release
|
|
193
|
-
uv pip install "astropy==7.2.0"
|
|
194
|
-
|
|
195
|
-
# Recommended optional dependencies for plotting and common workflows
|
|
196
|
-
uv pip install "astropy[recommended]==7.2.0"
|
|
197
|
-
|
|
198
|
-
# Full optional dependency set for broad astronomy workflows
|
|
199
|
-
uv pip install "astropy[all]==7.2.0"
|
|
200
|
-
```
|
|
201
|
-
|
|
202
|
-
Astropy 7.2.0 requires Python 3.11+ and depends on NumPy, PyERFA, PyYAML, and packaging. Use an isolated virtual environment; do not install Astropy with elevated privileges.
|
|
203
|
-
|
|
204
|
-
Note that the `[recommended]` and `[all]` extras pull in transitive dependencies (matplotlib, scipy, etc.) at unpinned versions. For reproducible production environments, pin the full dependency tree with a lockfile (`uv lock` in a project, or `uv pip compile` for requirements files) and review the resolved versions before deploying.
|
|
205
|
-
|
|
206
|
-
## Common Workflows
|
|
207
|
-
|
|
208
|
-
### Converting Coordinates Between Systems
|
|
209
|
-
|
|
210
|
-
```python
|
|
211
|
-
from astropy.coordinates import SkyCoord
|
|
212
|
-
import astropy.units as u
|
|
213
|
-
|
|
214
|
-
# Create coordinate
|
|
215
|
-
c = SkyCoord(ra='05h23m34.5s', dec='-69d45m22s', frame='icrs')
|
|
216
|
-
|
|
217
|
-
# Transform to galactic
|
|
218
|
-
c_gal = c.galactic
|
|
219
|
-
print(f"l={c_gal.l.deg}, b={c_gal.b.deg}")
|
|
220
|
-
|
|
221
|
-
# Transform to alt-az (requires time and location)
|
|
222
|
-
from astropy.time import Time
|
|
223
|
-
from astropy.coordinates import EarthLocation, AltAz
|
|
224
|
-
|
|
225
|
-
observing_time = Time('2023-06-15 23:00:00')
|
|
226
|
-
observing_location = EarthLocation(lat=40*u.deg, lon=-120*u.deg)
|
|
227
|
-
aa_frame = AltAz(obstime=observing_time, location=observing_location)
|
|
228
|
-
c_altaz = c.transform_to(aa_frame)
|
|
229
|
-
print(f"Alt={c_altaz.alt.deg}, Az={c_altaz.az.deg}")
|
|
230
|
-
```
|
|
231
|
-
|
|
232
|
-
### Reading and Analyzing FITS Files
|
|
233
|
-
|
|
234
|
-
```python
|
|
235
|
-
from astropy.io import fits
|
|
236
|
-
import numpy as np
|
|
237
|
-
|
|
238
|
-
# Open FITS file
|
|
239
|
-
with fits.open('observation.fits') as hdul:
|
|
240
|
-
# Display structure
|
|
241
|
-
hdul.info()
|
|
242
|
-
|
|
243
|
-
# Get image data and header
|
|
244
|
-
data = hdul[1].data
|
|
245
|
-
header = hdul[1].header
|
|
246
|
-
|
|
247
|
-
# Access header values
|
|
248
|
-
exptime = header['EXPTIME']
|
|
249
|
-
filter_name = header['FILTER']
|
|
250
|
-
|
|
251
|
-
# Analyze data
|
|
252
|
-
mean = np.mean(data)
|
|
253
|
-
median = np.median(data)
|
|
254
|
-
print(f"Mean: {mean}, Median: {median}")
|
|
255
|
-
```
|
|
256
|
-
|
|
257
|
-
### Cosmological Distance Calculations
|
|
258
|
-
|
|
259
|
-
```python
|
|
260
|
-
from astropy.cosmology import Planck18
|
|
261
|
-
import astropy.units as u
|
|
262
|
-
import numpy as np
|
|
263
|
-
|
|
264
|
-
# Calculate distances at z=1.5
|
|
265
|
-
z = 1.5
|
|
266
|
-
d_L = Planck18.luminosity_distance(z)
|
|
267
|
-
d_A = Planck18.angular_diameter_distance(z)
|
|
268
|
-
|
|
269
|
-
print(f"Luminosity distance: {d_L}")
|
|
270
|
-
print(f"Angular diameter distance: {d_A}")
|
|
271
|
-
|
|
272
|
-
# Age of universe at that redshift
|
|
273
|
-
age = Planck18.age(z)
|
|
274
|
-
print(f"Age at z={z}: {age.to(u.Gyr)}")
|
|
275
|
-
|
|
276
|
-
# Lookback time
|
|
277
|
-
t_lookback = Planck18.lookback_time(z)
|
|
278
|
-
print(f"Lookback time: {t_lookback.to(u.Gyr)}")
|
|
279
|
-
```
|
|
280
|
-
|
|
281
|
-
### Cross-Matching Catalogs
|
|
282
|
-
|
|
283
|
-
```python
|
|
284
|
-
from astropy.table import Table
|
|
285
|
-
from astropy.coordinates import SkyCoord, match_coordinates_sky
|
|
286
|
-
import astropy.units as u
|
|
287
|
-
|
|
288
|
-
# Read catalogs
|
|
289
|
-
cat1 = Table.read('catalog1.fits')
|
|
290
|
-
cat2 = Table.read('catalog2.fits')
|
|
291
|
-
|
|
292
|
-
# Create coordinate objects
|
|
293
|
-
coords1 = SkyCoord(ra=cat1['RA']*u.degree, dec=cat1['DEC']*u.degree)
|
|
294
|
-
coords2 = SkyCoord(ra=cat2['RA']*u.degree, dec=cat2['DEC']*u.degree)
|
|
295
|
-
|
|
296
|
-
# Find matches
|
|
297
|
-
idx, sep, _ = coords1.match_to_catalog_sky(coords2)
|
|
298
|
-
|
|
299
|
-
# Filter by separation threshold
|
|
300
|
-
max_sep = 1 * u.arcsec
|
|
301
|
-
matches = sep < max_sep
|
|
302
|
-
|
|
303
|
-
# Create matched catalogs
|
|
304
|
-
cat1_matched = cat1[matches]
|
|
305
|
-
cat2_matched = cat2[idx[matches]]
|
|
306
|
-
print(f"Found {len(cat1_matched)} matches")
|
|
307
|
-
```
|
|
308
|
-
|
|
309
|
-
## Best Practices
|
|
310
|
-
|
|
311
|
-
1. **Always use units**: Attach units to quantities to avoid errors and ensure dimensional consistency
|
|
312
|
-
2. **Use context managers for FITS files**: Ensures proper file closing
|
|
313
|
-
3. **Prefer arrays over loops**: Process multiple coordinates/times as arrays for better performance
|
|
314
|
-
4. **Check coordinate frames**: Verify the frame before transformations
|
|
315
|
-
5. **Use appropriate cosmology**: Choose the right cosmological model for your analysis
|
|
316
|
-
6. **Handle missing data**: Use masked columns for tables with missing values
|
|
317
|
-
7. **Specify time scales**: Be explicit about time scales (UTC, TT, TDB) for precise timing
|
|
318
|
-
8. **Use QTable for unit-aware tables**: When table columns have units
|
|
319
|
-
9. **Check WCS validity**: Verify WCS before using transformations
|
|
320
|
-
10. **Cache frequently used values**: Expensive calculations (e.g., cosmological distances) can be cached
|
|
321
|
-
11. **Be explicit about network access**: `SkyCoord.from_name()`, `EarthLocation.of_site(refresh_cache=True)`, `EarthLocation.of_address()`, `download_file()`, remote FITS reads, and some IERS time/coordinate transforms can contact external services or update local caches. Avoid sending sensitive target names, addresses, URLs, or proprietary file locations to third-party services. When working with potentially sensitive targets or data locations, confirm with the user before making these network calls.
|
|
322
|
-
12. **Pin for reproducibility**: Use pinned versions such as `astropy==7.2.0` for shared environments; update pins intentionally after reviewing release notes.
|
|
323
|
-
|
|
324
|
-
## Current-Version Notes
|
|
325
|
-
|
|
326
|
-
- Current stable release researched: Astropy 7.2.0 (released 2025-11-25; verified current as of 2026-06-10)
|
|
327
|
-
- Python requirement: 3.11+
|
|
328
|
-
- **Astropy 8.0 is at release-candidate stage** (8.0.0rc1, 2026-05-26). Key changes to anticipate:
|
|
329
|
-
- The deprecated `astropy.cosmology` submodule shims (`astropy.cosmology.flrw`, `.core`, `.funcs`, `.connect`, `.parameter`) are removed — import everything directly from `astropy.cosmology` (e.g., `from astropy.cosmology import FlatLambdaCDM, z_at_value`)
|
|
330
|
-
- `astropy.constants` defaults change from CODATA 2018 to CODATA 2022; pin a constants version via the `astropyconst` science states if reproducibility matters
|
|
331
|
-
- NumPy 2.0 becomes the minimum supported version; the 7.2.x LTS branch retains NumPy 1.x support for six months after the 8.0 release
|
|
332
|
-
- The built-in test runner (`astropy.test()`, `TestRunner`) is formally deprecated — invoke `pytest` directly
|
|
333
|
-
- Recent 7.x deprecations to avoid in new code: passing a table index identifier as the first `.loc` element (`t.loc["b", 2]`) — use `t.loc.with_index("b")[2]` instead (removal planned for 9.0); `astropy.utils.isiterable()` — use `numpy.iterable()`
|
|
334
|
-
- Recent 7.0 removals: older deprecated FITS APIs such as `(Bin)Table.update`, `_ExtensionHDU`, `_NonstandardExtHDU`, and the `tile_size` argument for `CompImageHDU`; `CompImageHeader` is deprecated. Avoid those legacy patterns in new examples.
|
|
335
|
-
- The recommended optional extras are `recommended` for common plotting/scientific dependencies and `all` only when a broad optional feature set is needed.
|
|
336
|
-
|
|
337
|
-
## Documentation and Resources
|
|
338
|
-
|
|
339
|
-
- Official Astropy Documentation: https://docs.astropy.org/en/stable/
|
|
340
|
-
- Tutorials: https://learn.astropy.org/
|
|
341
|
-
- GitHub: https://github.com/astropy/astropy
|
|
342
|
-
|
|
343
|
-
## Reference Files
|
|
344
|
-
|
|
345
|
-
For detailed information on specific modules:
|
|
346
|
-
- `references/units.md` - Units, quantities, conversions, and equivalencies
|
|
347
|
-
- `references/coordinates.md` - Coordinate systems, transformations, and catalog matching
|
|
348
|
-
- `references/cosmology.md` - Cosmological models and calculations
|
|
349
|
-
- `references/fits.md` - FITS file operations and manipulation
|
|
350
|
-
- `references/tables.md` - Table creation, I/O, and operations
|
|
351
|
-
- `references/time.md` - Time formats, scales, and calculations
|
|
352
|
-
- `references/wcs_and_other_modules.md` - WCS, NDData, modeling, visualization, constants, and utilities
|
|
353
|
-
|