@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,283 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: cellxgene-census
|
|
3
|
-
description: Query the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data. Use when you need population-scale cell metadata, gene expression slices, Census summary counts, source H5AD URIs/downloads, embeddings, spatial Census data, or reference atlas comparisons across organisms, tissues, diseases, assays, and cell types. For analyzing your own local single-cell data use scanpy, anndata, or scvi-tools.
|
|
4
|
-
allowed-tools: Read Write Edit Bash
|
|
5
|
-
license: MIT
|
|
6
|
-
compatibility: Requires Python >=3.10,<3.13. Examples target cellxgene-census 1.17.x and the 2025-11-08 stable LTS Census; spatial workflows need the spatial extra and TileDB-SOMA >=1.15.5. No authentication is required for public Census data.
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.2"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# CZ CELLxGENE Census
|
|
13
|
-
|
|
14
|
-
## Overview
|
|
15
|
-
|
|
16
|
-
The CZ CELLxGENE Census provides programmatic access to a comprehensive, versioned collection of standardized single-cell and spatial transcriptomics data from CZ CELLxGENE Discover. This skill enables efficient querying and analysis of public Census releases without downloading whole datasets first.
|
|
17
|
-
|
|
18
|
-
The Census includes:
|
|
19
|
-
- **217+ million total cells** and **125+ million unique cells** in the 2025-11-08 stable LTS release
|
|
20
|
-
- **1,845 datasets** in the 2025-11-08 stable LTS release
|
|
21
|
-
- **Human, mouse, marmoset, rhesus macaque, and chimpanzee** data in the current schema
|
|
22
|
-
- **Standardized metadata** (cell types, tissues, diseases, donors)
|
|
23
|
-
- **Raw gene expression** matrices and source H5AD lookup/download helpers
|
|
24
|
-
- **Pre-calculated summary counts, embeddings, and spatial data**
|
|
25
|
-
- **Integration with AnnData, Scanpy, TileDB-SOMA, TileDB-SOMA-ML, and other analysis tools**
|
|
26
|
-
|
|
27
|
-
## When to Use This Skill
|
|
28
|
-
|
|
29
|
-
This skill should be used when:
|
|
30
|
-
- Querying single-cell expression data by cell type, tissue, or disease
|
|
31
|
-
- Exploring available single-cell datasets and metadata
|
|
32
|
-
- Training machine learning models on single-cell data
|
|
33
|
-
- Performing large-scale cross-dataset analyses
|
|
34
|
-
- Integrating Census data with scanpy or other analysis frameworks
|
|
35
|
-
- Computing statistics across millions of cells
|
|
36
|
-
- Accessing pre-calculated embeddings or model predictions
|
|
37
|
-
|
|
38
|
-
## Installation and Setup
|
|
39
|
-
|
|
40
|
-
Install the Census API:
|
|
41
|
-
```bash
|
|
42
|
-
uv pip install "cellxgene-census==1.17.*"
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
For spatial workflows:
|
|
46
|
-
```bash
|
|
47
|
-
uv pip install "cellxgene-census[spatial]==1.17.*" "spatialdata[extra]>=0.2.5"
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
For PyTorch model training, use TileDB-SOMA-ML. The old `cellxgene_census.experimental.ml` loaders are deprecated:
|
|
51
|
-
|
|
52
|
-
```bash
|
|
53
|
-
uv pip install "cellxgene-census==1.17.*" tiledbsoma-ml
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
## Core Workflow Patterns
|
|
57
|
-
|
|
58
|
-
Eight patterns, each with code, are in
|
|
59
|
-
[references/core_workflow_patterns.md](references/core_workflow_patterns.md):
|
|
60
|
-
|
|
61
|
-
1. **Opening the Census** — always pin `census_version` so an analysis stays reproducible.
|
|
62
|
-
2. **Exploring Census information** — available datasets, cell counts, and summary tables.
|
|
63
|
-
3. **Querying expression data** — small to medium scale into an `AnnData`.
|
|
64
|
-
4. **Large-scale queries** — out-of-core processing when the slice will not fit in memory.
|
|
65
|
-
5. **Machine learning with PyTorch** — the Census data loaders.
|
|
66
|
-
6. **Spatial Census data** — accessing spatial assays.
|
|
67
|
-
7. **Integration with Scanpy** — handing a Census slice to a standard Scanpy workflow.
|
|
68
|
-
8. **Multi-dataset integration** — combining datasets and handling batch effects.
|
|
69
|
-
|
|
70
|
-
## Key Concepts and Best Practices
|
|
71
|
-
|
|
72
|
-
### Always Filter for Primary Data
|
|
73
|
-
Unless analyzing duplicates, always include `is_primary_data == True` in queries to avoid counting cells multiple times:
|
|
74
|
-
```python
|
|
75
|
-
obs_value_filter="cell_type == 'B cell' and is_primary_data == True"
|
|
76
|
-
```
|
|
77
|
-
|
|
78
|
-
### Specify Census Version for Reproducibility
|
|
79
|
-
Always specify the Census version in production analyses:
|
|
80
|
-
```python
|
|
81
|
-
census = cellxgene_census.open_soma(census_version="2025-11-08")
|
|
82
|
-
```
|
|
83
|
-
|
|
84
|
-
### Estimate Query Size Before Loading
|
|
85
|
-
For large queries, first check the number of cells to avoid memory issues:
|
|
86
|
-
```python
|
|
87
|
-
# Get cell count
|
|
88
|
-
metadata = cellxgene_census.get_obs(
|
|
89
|
-
census, "homo_sapiens",
|
|
90
|
-
value_filter="tissue_general == 'brain' and is_primary_data == True",
|
|
91
|
-
column_names=["soma_joinid"]
|
|
92
|
-
)
|
|
93
|
-
n_cells = len(metadata)
|
|
94
|
-
print(f"Query will return {n_cells:,} cells")
|
|
95
|
-
|
|
96
|
-
# If too large (>100k), use out-of-core processing
|
|
97
|
-
```
|
|
98
|
-
|
|
99
|
-
### Use tissue_general for Broader Groupings
|
|
100
|
-
The `tissue_general` field provides coarser categories than `tissue`, useful for cross-tissue analyses:
|
|
101
|
-
```python
|
|
102
|
-
# Broader grouping
|
|
103
|
-
obs_value_filter="tissue_general == 'immune system'"
|
|
104
|
-
|
|
105
|
-
# Specific tissue
|
|
106
|
-
obs_value_filter="tissue == 'peripheral blood mononuclear cell'"
|
|
107
|
-
```
|
|
108
|
-
|
|
109
|
-
### Select Only Needed Columns
|
|
110
|
-
Minimize data transfer by specifying only required metadata columns:
|
|
111
|
-
```python
|
|
112
|
-
obs_column_names=["cell_type", "tissue_general", "disease"] # Not all columns
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
### Check Dataset Presence for Gene-Specific Queries
|
|
116
|
-
When analyzing specific genes, verify which datasets measured them:
|
|
117
|
-
```python
|
|
118
|
-
presence = cellxgene_census.get_presence_matrix(
|
|
119
|
-
census,
|
|
120
|
-
"homo_sapiens",
|
|
121
|
-
var_value_filter="feature_name in ['CD4', 'CD8A']"
|
|
122
|
-
)
|
|
123
|
-
```
|
|
124
|
-
|
|
125
|
-
### Two-Step Workflow: Explore Then Query
|
|
126
|
-
First explore metadata to understand available data, then query expression:
|
|
127
|
-
```python
|
|
128
|
-
# Step 1: Explore what's available
|
|
129
|
-
metadata = cellxgene_census.get_obs(
|
|
130
|
-
census, "homo_sapiens",
|
|
131
|
-
value_filter="disease == 'COVID-19' and is_primary_data == True",
|
|
132
|
-
column_names=["cell_type", "tissue_general"]
|
|
133
|
-
)
|
|
134
|
-
print(metadata.value_counts())
|
|
135
|
-
|
|
136
|
-
# Step 2: Query based on findings
|
|
137
|
-
adata = cellxgene_census.get_anndata(
|
|
138
|
-
census=census,
|
|
139
|
-
organism="Homo sapiens",
|
|
140
|
-
obs_value_filter="disease == 'COVID-19' and cell_type == 'T cell' and is_primary_data == True",
|
|
141
|
-
)
|
|
142
|
-
```
|
|
143
|
-
|
|
144
|
-
## Available Metadata Fields
|
|
145
|
-
|
|
146
|
-
### Cell Metadata (obs)
|
|
147
|
-
Key fields for filtering:
|
|
148
|
-
- `cell_type`, `cell_type_ontology_term_id`
|
|
149
|
-
- `tissue`, `tissue_general`, `tissue_ontology_term_id`
|
|
150
|
-
- `disease`, `disease_ontology_term_id`
|
|
151
|
-
- `assay`, `assay_ontology_term_id`
|
|
152
|
-
- `donor_id`, `sex`, `self_reported_ethnicity`
|
|
153
|
-
- `development_stage`, `development_stage_ontology_term_id`
|
|
154
|
-
- `dataset_id`
|
|
155
|
-
- `is_primary_data` (Boolean: True = unique cell)
|
|
156
|
-
|
|
157
|
-
The current schema includes organism collections beyond human and mouse. Confirm available organisms for the selected release with `list(census["census_data"].keys())`.
|
|
158
|
-
|
|
159
|
-
### Gene Metadata (var)
|
|
160
|
-
- `feature_id` (Ensembl gene ID, e.g., "ENSG00000161798")
|
|
161
|
-
- `feature_name` (Gene symbol, e.g., "FOXP2")
|
|
162
|
-
- `feature_type`
|
|
163
|
-
- `feature_length` (Gene length in base pairs)
|
|
164
|
-
- `nnz`, `n_measured_obs` (availability summaries useful for checking sparsity and coverage)
|
|
165
|
-
|
|
166
|
-
## Reference Documentation
|
|
167
|
-
|
|
168
|
-
This skill includes detailed reference documentation:
|
|
169
|
-
|
|
170
|
-
### references/census_schema.md
|
|
171
|
-
Comprehensive documentation of:
|
|
172
|
-
- Census data structure and organization
|
|
173
|
-
- All available metadata fields
|
|
174
|
-
- Value filter syntax and operators
|
|
175
|
-
- SOMA object types
|
|
176
|
-
- Data inclusion criteria
|
|
177
|
-
|
|
178
|
-
**When to read:** When you need detailed schema information, full list of metadata fields, or complex filter syntax.
|
|
179
|
-
|
|
180
|
-
### references/common_patterns.md
|
|
181
|
-
Examples and patterns for:
|
|
182
|
-
- Exploratory queries (metadata only)
|
|
183
|
-
- Small-to-medium queries (AnnData)
|
|
184
|
-
- Large queries (out-of-core processing)
|
|
185
|
-
- PyTorch integration
|
|
186
|
-
- Spatial Census access patterns
|
|
187
|
-
- Scanpy integration workflows
|
|
188
|
-
- Multi-dataset integration
|
|
189
|
-
- Best practices and common pitfalls
|
|
190
|
-
|
|
191
|
-
**When to read:** When implementing specific query patterns, looking for code examples, or troubleshooting common issues.
|
|
192
|
-
|
|
193
|
-
## Common Use Cases
|
|
194
|
-
|
|
195
|
-
### Use Case 1: Explore Cell Types in a Tissue
|
|
196
|
-
```python
|
|
197
|
-
with cellxgene_census.open_soma() as census:
|
|
198
|
-
cells = cellxgene_census.get_obs(
|
|
199
|
-
census, "homo_sapiens",
|
|
200
|
-
value_filter="tissue_general == 'lung' and is_primary_data == True",
|
|
201
|
-
column_names=["cell_type"]
|
|
202
|
-
)
|
|
203
|
-
print(cells["cell_type"].value_counts())
|
|
204
|
-
```
|
|
205
|
-
|
|
206
|
-
### Use Case 2: Query Marker Gene Expression
|
|
207
|
-
```python
|
|
208
|
-
with cellxgene_census.open_soma() as census:
|
|
209
|
-
adata = cellxgene_census.get_anndata(
|
|
210
|
-
census=census,
|
|
211
|
-
organism="Homo sapiens",
|
|
212
|
-
var_value_filter="feature_name in ['CD4', 'CD8A', 'CD19']",
|
|
213
|
-
obs_value_filter="cell_type in ['T cell', 'B cell'] and is_primary_data == True",
|
|
214
|
-
)
|
|
215
|
-
```
|
|
216
|
-
|
|
217
|
-
### Use Case 3: Train Cell Type Classifier
|
|
218
|
-
```python
|
|
219
|
-
import tiledbsoma as soma
|
|
220
|
-
from tiledbsoma_ml import ExperimentDataset, experiment_dataloader
|
|
221
|
-
|
|
222
|
-
with cellxgene_census.open_soma() as census:
|
|
223
|
-
experiment = census["census_data"]["homo_sapiens"]
|
|
224
|
-
with experiment.axis_query(
|
|
225
|
-
measurement_name="RNA",
|
|
226
|
-
obs_query=soma.AxisQuery(value_filter="is_primary_data == True"),
|
|
227
|
-
) as query:
|
|
228
|
-
dataset = ExperimentDataset(
|
|
229
|
-
query=query,
|
|
230
|
-
layer_name="raw",
|
|
231
|
-
obs_column_names=["cell_type"],
|
|
232
|
-
batch_size=128,
|
|
233
|
-
shuffle=True,
|
|
234
|
-
)
|
|
235
|
-
dataloader = experiment_dataloader(dataset)
|
|
236
|
-
|
|
237
|
-
for X, obs in dataloader:
|
|
238
|
-
labels = obs["cell_type"]
|
|
239
|
-
# Training logic
|
|
240
|
-
pass
|
|
241
|
-
```
|
|
242
|
-
|
|
243
|
-
### Use Case 4: Cross-Tissue Analysis
|
|
244
|
-
```python
|
|
245
|
-
with cellxgene_census.open_soma() as census:
|
|
246
|
-
adata = cellxgene_census.get_anndata(
|
|
247
|
-
census=census,
|
|
248
|
-
organism="Homo sapiens",
|
|
249
|
-
obs_value_filter="cell_type == 'macrophage' and tissue_general in ['lung', 'liver', 'brain'] and is_primary_data == True",
|
|
250
|
-
)
|
|
251
|
-
|
|
252
|
-
# Analyze macrophage differences across tissues
|
|
253
|
-
sc.tl.rank_genes_groups(adata, groupby="tissue_general")
|
|
254
|
-
```
|
|
255
|
-
|
|
256
|
-
## Troubleshooting
|
|
257
|
-
|
|
258
|
-
### Query Returns Too Many Cells
|
|
259
|
-
- Add more specific filters to reduce scope
|
|
260
|
-
- Use `tissue` instead of `tissue_general` for finer granularity
|
|
261
|
-
- Filter by specific `dataset_id` if known
|
|
262
|
-
- Switch to out-of-core processing for large queries
|
|
263
|
-
|
|
264
|
-
### Memory Errors
|
|
265
|
-
- Reduce query scope with more restrictive filters
|
|
266
|
-
- Select fewer genes with `var_value_filter`
|
|
267
|
-
- Use out-of-core processing with `axis_query()`
|
|
268
|
-
- Process data in batches
|
|
269
|
-
|
|
270
|
-
### Duplicate Cells in Results
|
|
271
|
-
- Always include `is_primary_data == True` in filters
|
|
272
|
-
- Check if intentionally querying across multiple datasets
|
|
273
|
-
|
|
274
|
-
### Gene Not Found
|
|
275
|
-
- Verify gene name spelling (case-sensitive)
|
|
276
|
-
- Try Ensembl ID with `feature_id` instead of `feature_name`
|
|
277
|
-
- Check dataset presence matrix to see if gene was measured
|
|
278
|
-
- Some genes may have been filtered during Census construction
|
|
279
|
-
|
|
280
|
-
### Version Inconsistencies
|
|
281
|
-
- Always specify `census_version` explicitly
|
|
282
|
-
- Use same version across all analyses
|
|
283
|
-
- Check release notes for version-specific changes
|
package/skills/cirq/SKILL.md
DELETED
|
@@ -1,370 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: cirq
|
|
3
|
-
description: Google quantum computing framework. Use when targeting Google Quantum AI hardware, designing noise-aware circuits, or running quantum characterization experiments. Best for Google hardware, noise modeling, and low-level circuit design. For IBM hardware use qiskit; for quantum ML with autodiff use pennylane; for physics simulations use qutip.
|
|
4
|
-
license: Apache-2.0 license
|
|
5
|
-
allowed-tools: Read Write Edit Bash
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.0"
|
|
8
|
-
skill-author: K-Dense Inc.
|
|
9
|
-
---
|
|
10
|
-
|
|
11
|
-
# Cirq - Quantum Computing with Python
|
|
12
|
-
|
|
13
|
-
Cirq is Google Quantum AI's open-source framework for designing, simulating, and running quantum circuits on quantum computers and simulators.
|
|
14
|
-
|
|
15
|
-
## When to Use This Skill
|
|
16
|
-
|
|
17
|
-
Use this skill when:
|
|
18
|
-
- Building, simulating, or optimizing NISQ circuits in Python
|
|
19
|
-
- Running jobs on Google Quantum AI processors (via `cirq-google`) or partner backends (IonQ, Azure Quantum, AQT, Pasqal)
|
|
20
|
-
- Modeling noise, compiling to hardware gatesets, or designing characterization experiments
|
|
21
|
-
- Using parameter sweeps, transformers, or the ReCirq experiment patterns
|
|
22
|
-
|
|
23
|
-
For IBM hardware use **qiskit**; for quantum ML with autodiff use **pennylane**; for physics simulations use **qutip**.
|
|
24
|
-
|
|
25
|
-
## Installation
|
|
26
|
-
|
|
27
|
-
Requires Python 3.11+. Current stable release: **1.6.1** (August 2025). Vendor packages share the same version number.
|
|
28
|
-
|
|
29
|
-
```bash
|
|
30
|
-
uv pip install "cirq==1.6.1"
|
|
31
|
-
```
|
|
32
|
-
|
|
33
|
-
For hardware integration (pin matching versions for reproducibility):
|
|
34
|
-
```bash
|
|
35
|
-
# Google Quantum Engine (requires approved GCP project access)
|
|
36
|
-
uv pip install "cirq-google==1.6.1"
|
|
37
|
-
|
|
38
|
-
# IonQ
|
|
39
|
-
uv pip install "cirq-ionq==1.6.1"
|
|
40
|
-
|
|
41
|
-
# AQT (Alpine Quantum Technologies)
|
|
42
|
-
uv pip install "cirq-aqt==1.6.1"
|
|
43
|
-
|
|
44
|
-
# Pasqal
|
|
45
|
-
uv pip install "cirq-pasqal==1.6.1"
|
|
46
|
-
|
|
47
|
-
# Azure Quantum (IonQ, Honeywell/Quantinuum backends)
|
|
48
|
-
uv pip install "azure-quantum[cirq]"
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
For latest features during development, omit version pins; for production or hardware runs, pin all packages to the same Cirq release.
|
|
52
|
-
|
|
53
|
-
## Quick Start
|
|
54
|
-
|
|
55
|
-
### Basic Circuit
|
|
56
|
-
|
|
57
|
-
```python
|
|
58
|
-
import cirq
|
|
59
|
-
import numpy as np
|
|
60
|
-
|
|
61
|
-
# Create qubits
|
|
62
|
-
q0, q1 = cirq.LineQubit.range(2)
|
|
63
|
-
|
|
64
|
-
# Build circuit
|
|
65
|
-
circuit = cirq.Circuit(
|
|
66
|
-
cirq.H(q0), # Hadamard on q0
|
|
67
|
-
cirq.CNOT(q0, q1), # CNOT with q0 control, q1 target
|
|
68
|
-
cirq.measure(q0, q1, key='result')
|
|
69
|
-
)
|
|
70
|
-
|
|
71
|
-
print(circuit)
|
|
72
|
-
|
|
73
|
-
# Simulate
|
|
74
|
-
simulator = cirq.Simulator()
|
|
75
|
-
result = simulator.run(circuit, repetitions=1000)
|
|
76
|
-
|
|
77
|
-
# Display results
|
|
78
|
-
print(result.histogram(key='result'))
|
|
79
|
-
```
|
|
80
|
-
|
|
81
|
-
### Parameterized Circuit
|
|
82
|
-
|
|
83
|
-
```python
|
|
84
|
-
import sympy
|
|
85
|
-
|
|
86
|
-
# Define symbolic parameter
|
|
87
|
-
theta = sympy.Symbol('theta')
|
|
88
|
-
|
|
89
|
-
# Create parameterized circuit
|
|
90
|
-
circuit = cirq.Circuit(
|
|
91
|
-
cirq.ry(theta)(q0),
|
|
92
|
-
cirq.measure(q0, key='m')
|
|
93
|
-
)
|
|
94
|
-
|
|
95
|
-
# Sweep over parameter values
|
|
96
|
-
sweep = cirq.Linspace('theta', start=0, stop=2*np.pi, length=20)
|
|
97
|
-
results = simulator.run_sweep(circuit, params=sweep, repetitions=1000)
|
|
98
|
-
|
|
99
|
-
# Process results
|
|
100
|
-
for params, result in zip(sweep, results):
|
|
101
|
-
theta_val = params['theta']
|
|
102
|
-
counts = result.histogram(key='m')
|
|
103
|
-
print(f"θ={theta_val:.2f}: {counts}")
|
|
104
|
-
```
|
|
105
|
-
|
|
106
|
-
## Core Capabilities
|
|
107
|
-
|
|
108
|
-
### Circuit Building
|
|
109
|
-
For comprehensive information about building quantum circuits, including qubits, gates, operations, custom gates, and circuit patterns, see:
|
|
110
|
-
- **[references/building.md](references/building.md)** - Complete guide to circuit construction
|
|
111
|
-
|
|
112
|
-
Common topics:
|
|
113
|
-
- Qubit types (GridQubit, LineQubit, NamedQubit)
|
|
114
|
-
- Single and two-qubit gates
|
|
115
|
-
- Parameterized gates and operations
|
|
116
|
-
- Custom gate decomposition
|
|
117
|
-
- Circuit organization with moments
|
|
118
|
-
- Standard circuit patterns (Bell states, GHZ, QFT)
|
|
119
|
-
- Import/export (OpenQASM, JSON)
|
|
120
|
-
- Working with qudits and observables
|
|
121
|
-
|
|
122
|
-
### Simulation
|
|
123
|
-
For detailed information about simulating quantum circuits, including exact simulation, noisy simulation, parameter sweeps, and the Quantum Virtual Machine, see:
|
|
124
|
-
- **[references/simulation.md](references/simulation.md)** - Complete guide to quantum simulation
|
|
125
|
-
|
|
126
|
-
Common topics:
|
|
127
|
-
- Exact simulation (state vector, density matrix)
|
|
128
|
-
- Sampling and measurements
|
|
129
|
-
- Parameter sweeps (single and multiple parameters)
|
|
130
|
-
- Noisy simulation
|
|
131
|
-
- State histograms and visualization
|
|
132
|
-
- Quantum Virtual Machine (QVM)
|
|
133
|
-
- Expectation values and observables
|
|
134
|
-
- Performance optimization
|
|
135
|
-
|
|
136
|
-
### Circuit Transformation
|
|
137
|
-
For information about optimizing, compiling, and manipulating quantum circuits, see:
|
|
138
|
-
- **[references/transformation.md](references/transformation.md)** - Complete guide to circuit transformations
|
|
139
|
-
|
|
140
|
-
Common topics:
|
|
141
|
-
- Transformer framework
|
|
142
|
-
- Gate decomposition
|
|
143
|
-
- Circuit optimization (merge gates, eject Z gates, drop negligible operations)
|
|
144
|
-
- Circuit compilation for hardware
|
|
145
|
-
- Qubit routing and SWAP insertion
|
|
146
|
-
- Custom transformers
|
|
147
|
-
- Transformation pipelines
|
|
148
|
-
|
|
149
|
-
### Hardware Integration
|
|
150
|
-
For information about running circuits on real quantum hardware from various providers, see:
|
|
151
|
-
- **[references/hardware.md](references/hardware.md)** - Complete guide to hardware integration
|
|
152
|
-
|
|
153
|
-
Supported providers:
|
|
154
|
-
- **Google Quantum AI** (`cirq-google`) — Sycamore, Weber, Willow processors via Quantum Engine (restricted access; requires approved GCP project)
|
|
155
|
-
- **IonQ** (`cirq-ionq`) — trapped-ion QPUs and simulators
|
|
156
|
-
- **Azure Quantum** (`azure-quantum[cirq]`) — IonQ and Honeywell/Quantinuum backends
|
|
157
|
-
- **AQT** (`cirq-aqt`) — Alpine Quantum Technologies
|
|
158
|
-
- **Pasqal** (`cirq-pasqal`) — neutral-atom devices
|
|
159
|
-
|
|
160
|
-
Topics include device representation, qubit selection, authentication, job management, and circuit optimization for hardware. See [Access and authentication](https://quantumai.google/cirq/google/access) for Google Cloud setup.
|
|
161
|
-
|
|
162
|
-
### Noise Modeling
|
|
163
|
-
For information about modeling noise, noisy simulation, characterization, and error mitigation, see:
|
|
164
|
-
- **[references/noise.md](references/noise.md)** - Complete guide to noise modeling
|
|
165
|
-
|
|
166
|
-
Common topics:
|
|
167
|
-
- Noise channels (depolarizing, amplitude damping, phase damping)
|
|
168
|
-
- Noise models (constant, gate-specific, qubit-specific, thermal)
|
|
169
|
-
- Adding noise to circuits
|
|
170
|
-
- Readout noise
|
|
171
|
-
- Noise characterization (randomized benchmarking, XEB)
|
|
172
|
-
- Noise visualization (heatmaps)
|
|
173
|
-
- Error mitigation techniques
|
|
174
|
-
|
|
175
|
-
### Quantum Experiments
|
|
176
|
-
For information about designing experiments, parameter sweeps, data collection, and using the ReCirq framework, see:
|
|
177
|
-
- **[references/experiments.md](references/experiments.md)** - Complete guide to quantum experiments
|
|
178
|
-
|
|
179
|
-
Common topics:
|
|
180
|
-
- Experiment design patterns
|
|
181
|
-
- Parameter sweeps and data collection
|
|
182
|
-
- ReCirq framework structure
|
|
183
|
-
- Common algorithms (VQE, QAOA, QPE)
|
|
184
|
-
- Data analysis and visualization
|
|
185
|
-
- Statistical analysis and fidelity estimation
|
|
186
|
-
- Parallel data collection
|
|
187
|
-
|
|
188
|
-
## Common Patterns
|
|
189
|
-
|
|
190
|
-
### Variational Algorithm Template
|
|
191
|
-
|
|
192
|
-
```python
|
|
193
|
-
import scipy.optimize
|
|
194
|
-
|
|
195
|
-
def variational_algorithm(ansatz, cost_function, initial_params):
|
|
196
|
-
"""Template for variational quantum algorithms."""
|
|
197
|
-
|
|
198
|
-
def objective(params):
|
|
199
|
-
circuit = ansatz(params)
|
|
200
|
-
simulator = cirq.Simulator()
|
|
201
|
-
result = simulator.simulate(circuit)
|
|
202
|
-
return cost_function(result)
|
|
203
|
-
|
|
204
|
-
# Optimize
|
|
205
|
-
result = scipy.optimize.minimize(
|
|
206
|
-
objective,
|
|
207
|
-
initial_params,
|
|
208
|
-
method='COBYLA'
|
|
209
|
-
)
|
|
210
|
-
|
|
211
|
-
return result
|
|
212
|
-
|
|
213
|
-
# Define ansatz
|
|
214
|
-
def my_ansatz(params):
|
|
215
|
-
q = cirq.LineQubit(0)
|
|
216
|
-
return cirq.Circuit(
|
|
217
|
-
cirq.ry(params[0])(q),
|
|
218
|
-
cirq.rz(params[1])(q)
|
|
219
|
-
)
|
|
220
|
-
|
|
221
|
-
# Define cost function
|
|
222
|
-
def my_cost(result):
|
|
223
|
-
state = result.final_state_vector
|
|
224
|
-
# Calculate cost based on state
|
|
225
|
-
return np.real(state[0])
|
|
226
|
-
|
|
227
|
-
# Run optimization
|
|
228
|
-
result = variational_algorithm(my_ansatz, my_cost, [0.0, 0.0])
|
|
229
|
-
```
|
|
230
|
-
|
|
231
|
-
### Hardware Execution Template
|
|
232
|
-
|
|
233
|
-
```python
|
|
234
|
-
import os
|
|
235
|
-
|
|
236
|
-
def run_on_hardware(circuit, provider='google', processor_id=None, repetitions=1000):
|
|
237
|
-
"""Template for running on quantum hardware."""
|
|
238
|
-
|
|
239
|
-
if provider == 'google':
|
|
240
|
-
import cirq_google as cg
|
|
241
|
-
|
|
242
|
-
project_id = os.environ['GOOGLE_CLOUD_PROJECT']
|
|
243
|
-
engine = cg.Engine(project_id=project_id)
|
|
244
|
-
|
|
245
|
-
# List available processors: engine.list_processors()
|
|
246
|
-
processor_id = processor_id or 'weber' # use your assigned processor_id
|
|
247
|
-
sampler = engine.get_sampler(processor_id=processor_id)
|
|
248
|
-
return sampler.run(circuit, repetitions=repetitions)
|
|
249
|
-
|
|
250
|
-
elif provider == 'ionq':
|
|
251
|
-
import cirq_ionq as ionq
|
|
252
|
-
|
|
253
|
-
# Requires IONQ_API_KEY in environment
|
|
254
|
-
service = ionq.Service()
|
|
255
|
-
return service.run(circuit, repetitions=repetitions, target='qpu')
|
|
256
|
-
|
|
257
|
-
elif provider == 'azure':
|
|
258
|
-
from azure.quantum.cirq import AzureQuantumService
|
|
259
|
-
|
|
260
|
-
service = AzureQuantumService(
|
|
261
|
-
resource_id=os.environ['AZURE_QUANTUM_RESOURCE_ID'],
|
|
262
|
-
location=os.environ['AZURE_QUANTUM_LOCATION'],
|
|
263
|
-
)
|
|
264
|
-
return service.run(circuit, repetitions=repetitions, target='ionq.qpu')
|
|
265
|
-
|
|
266
|
-
else:
|
|
267
|
-
raise ValueError(f"Unknown provider: {provider}")
|
|
268
|
-
```
|
|
269
|
-
|
|
270
|
-
### Noise Study Template
|
|
271
|
-
|
|
272
|
-
```python
|
|
273
|
-
def noise_comparison_study(circuit, noise_levels):
|
|
274
|
-
"""Compare circuit performance at different noise levels."""
|
|
275
|
-
|
|
276
|
-
results = {}
|
|
277
|
-
|
|
278
|
-
for noise_level in noise_levels:
|
|
279
|
-
# Create noisy circuit
|
|
280
|
-
noisy_circuit = circuit.with_noise(cirq.depolarize(p=noise_level))
|
|
281
|
-
|
|
282
|
-
# Simulate
|
|
283
|
-
simulator = cirq.DensityMatrixSimulator()
|
|
284
|
-
result = simulator.run(noisy_circuit, repetitions=1000)
|
|
285
|
-
|
|
286
|
-
# Analyze
|
|
287
|
-
results[noise_level] = {
|
|
288
|
-
'histogram': result.histogram(key='result'),
|
|
289
|
-
'dominant_state': max(
|
|
290
|
-
result.histogram(key='result').items(),
|
|
291
|
-
key=lambda x: x[1]
|
|
292
|
-
)
|
|
293
|
-
}
|
|
294
|
-
|
|
295
|
-
return results
|
|
296
|
-
|
|
297
|
-
# Run study
|
|
298
|
-
noise_levels = [0.0, 0.001, 0.01, 0.05, 0.1]
|
|
299
|
-
results = noise_comparison_study(circuit, noise_levels)
|
|
300
|
-
```
|
|
301
|
-
|
|
302
|
-
## Best Practices
|
|
303
|
-
|
|
304
|
-
1. **Circuit Design**
|
|
305
|
-
- Use appropriate qubit types for your topology
|
|
306
|
-
- Keep circuits modular and reusable
|
|
307
|
-
- Label measurements with descriptive keys
|
|
308
|
-
- Validate circuits against device constraints before execution
|
|
309
|
-
|
|
310
|
-
2. **Simulation**
|
|
311
|
-
- Use state vector simulation for pure states (more efficient)
|
|
312
|
-
- Use density matrix simulation only when needed (mixed states, noise)
|
|
313
|
-
- Leverage parameter sweeps instead of individual runs
|
|
314
|
-
- Monitor memory usage for large systems (2^n grows quickly)
|
|
315
|
-
|
|
316
|
-
3. **Hardware Execution**
|
|
317
|
-
- Always test on simulators first
|
|
318
|
-
- Select best qubits using calibration data
|
|
319
|
-
- Optimize circuits for target hardware gateset
|
|
320
|
-
- Implement error mitigation for production runs
|
|
321
|
-
- Store expensive hardware results immediately
|
|
322
|
-
|
|
323
|
-
4. **Circuit Optimization**
|
|
324
|
-
- Start with high-level built-in transformers
|
|
325
|
-
- Chain multiple optimizations in sequence
|
|
326
|
-
- Track depth and gate count reduction
|
|
327
|
-
- Validate correctness after transformation
|
|
328
|
-
|
|
329
|
-
5. **Noise Modeling**
|
|
330
|
-
- Use realistic noise models from calibration data
|
|
331
|
-
- Include all error sources (gate, decoherence, readout)
|
|
332
|
-
- Characterize before mitigating
|
|
333
|
-
- Keep circuits shallow to minimize noise accumulation
|
|
334
|
-
|
|
335
|
-
6. **Experiments**
|
|
336
|
-
- Structure experiments with clear separation (data generation, collection, analysis)
|
|
337
|
-
- Use ReCirq patterns for reproducibility
|
|
338
|
-
- Save intermediate results frequently
|
|
339
|
-
- Parallelize independent tasks
|
|
340
|
-
- Document thoroughly with metadata
|
|
341
|
-
|
|
342
|
-
## Additional Resources
|
|
343
|
-
|
|
344
|
-
- **Official Documentation**: https://quantumai.google/cirq
|
|
345
|
-
- **API Reference**: https://quantumai.google/reference/python/cirq
|
|
346
|
-
- **Tutorials**: https://quantumai.google/cirq/tutorials
|
|
347
|
-
- **Examples**: https://github.com/quantumlib/Cirq/tree/main/examples
|
|
348
|
-
- **Version policy**: https://quantumai.google/cirq/dev/versions
|
|
349
|
-
- **ReCirq**: https://github.com/quantumlib/ReCirq
|
|
350
|
-
|
|
351
|
-
## Common Issues
|
|
352
|
-
|
|
353
|
-
**Circuit too deep for hardware:**
|
|
354
|
-
- Use circuit optimization transformers to reduce depth
|
|
355
|
-
- See `transformation.md` for optimization techniques
|
|
356
|
-
|
|
357
|
-
**Memory issues with simulation:**
|
|
358
|
-
- Switch from density matrix to state vector simulator
|
|
359
|
-
- Reduce number of qubits or use stabilizer simulator for Clifford circuits
|
|
360
|
-
|
|
361
|
-
**Device validation errors:**
|
|
362
|
-
- Check qubit connectivity with device.metadata.nx_graph
|
|
363
|
-
- Decompose gates to device-native gateset
|
|
364
|
-
- See `hardware.md` for device-specific compilation
|
|
365
|
-
|
|
366
|
-
**Noisy simulation too slow:**
|
|
367
|
-
- Density matrix simulation is O(2^2n) - consider reducing qubits
|
|
368
|
-
- Use noise models selectively on critical operations only
|
|
369
|
-
- See `simulation.md` for performance optimization
|
|
370
|
-
|