@pikaa-ai/pikaa 0.3.23 → 0.3.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +407 -219
  6. package/dist/index.js +7 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,303 +0,0 @@
1
- ---
2
- name: scanpy
3
- description: Standard single-cell RNA-seq analysis pipeline. Use for QC, normalization, dimensionality reduction (PCA/UMAP/t-SNE), clustering, differential expression, visualization, and converting R-friendly single-cell formats such as Seurat or SingleCellExperiment RDS files into h5ad for Scanpy. Best for exploratory scRNA-seq analysis with established workflows. For deep learning models use scvi-tools; for data format questions use anndata.
4
- license: BSD-3-Clause
5
- metadata:
6
- version: "1.5"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # Scanpy: Single-Cell Analysis
11
-
12
- ## Overview
13
-
14
- Scanpy is a scalable Python toolkit for analyzing single-cell RNA-seq data, built on AnnData. Apply this skill for complete single-cell workflows including quality control, normalization, dimensionality reduction, clustering, marker gene identification, visualization, and trajectory analysis. Current stable release: **scanpy 1.12.x** (January 2026).
15
-
16
- ## Installation
17
-
18
- Requires Python **3.12+** (scanpy 1.12 dropped Python ≤3.11) and anndata **≥0.10**.
19
-
20
- ```bash
21
- uv pip install "scanpy[leiden]"
22
- ```
23
-
24
- The `[leiden]` extra installs `python-igraph` and `leidenalg`, required for Leiden clustering. For reproducible environments, pin a version: `uv pip install "scanpy[leiden]==1.12.1"`.
25
-
26
- For large or out-of-core datasets, many functions support [Dask](https://docs.dask.org/) arrays (experimental):
27
-
28
- ```bash
29
- uv pip install "scanpy[leiden]" dask
30
- ```
31
-
32
- See the [Using dask with Scanpy](https://scanpy.scverse.org/en/stable/tutorials/experimental/dask.html) tutorial. For GPU-accelerated scanpy-like operations, use [rapids-singlecell](https://rapids-singlecell.readthedocs.io/) as a separate package.
33
-
34
- If the input is an R-native single-cell object (`.rds`, `.RData`, Seurat, or SingleCellExperiment), first convert it to `.h5ad` with R tooling, then load it with Scanpy. Read `references/r_interop.md` for agent-run installation and conversion instructions across macOS, Linux, and Windows.
35
-
36
- For AnnData structure and I/O details, use the **anndata** skill. For probabilistic models and batch correction, use **scvi-tools**.
37
-
38
- ## When to Use This Skill
39
-
40
- This skill should be used when:
41
- - Analyzing single-cell RNA-seq data (.h5ad, 10X, CSV formats)
42
- - Working with R-friendly single-cell datasets (`.rds`, `.RData`, Seurat, SingleCellExperiment) that need conversion to `.h5ad`
43
- - Performing quality control on scRNA-seq datasets
44
- - Creating UMAP, t-SNE, or PCA visualizations
45
- - Identifying cell clusters and finding marker genes
46
- - Annotating cell types based on gene expression
47
- - Conducting trajectory inference or pseudotime analysis
48
- - Generating publication-quality single-cell plots
49
-
50
- ## Script Toolkit (prefer these over writing code from scratch)
51
-
52
- This skill bundles ready-to-run CLI scripts in `scripts/` for every common step. **Run these instead of hand-writing scanpy code** — they handle file loading by extension, figure setup, sensible defaults, raw-count preservation, and progress logging. Each reads and writes `.h5ad`, so they chain together, and each has its own `--help`. Only drop down to writing scanpy code when a task isn't covered by a script or needs unusual customization.
53
-
54
- All scripts use a shared `scripts/_common.py` helper (loading, saving, figure config) — keep it alongside the others. Run from the skill directory or pass full paths; figures default to `./figures/`.
55
-
56
- | Script | Purpose | Typical call |
57
- |--------|---------|--------------|
58
- | `run_pipeline.py` | **Full workflow in one command**: load → QC → normalize → HVG → PCA → (batch) → UMAP → Leiden → markers | `python scripts/run_pipeline.py raw.h5ad -o processed.h5ad` |
59
- | `inspect_data.py` | Summarize an unknown dataset (shape, obs/var, layers, what's already computed, raw vs normalized) | `python scripts/inspect_data.py data.h5ad` |
60
- | `convert.py` | Load any format (10x dir/.h5, csv, loom, mtx) and write `.h5ad` | `python scripts/convert.py 10x_dir/ -o data.h5ad` |
61
- | `qc_analysis.py` | QC metrics, before/after plots, filtering, optional Scrublet doublets | `python scripts/qc_analysis.py raw.h5ad -o qc.h5ad --scrublet` |
62
- | `preprocess.py` | Normalize, log1p, HVG, optional scale/regress (keeps `counts` layer + `raw`) | `python scripts/preprocess.py qc.h5ad -o norm.h5ad` |
63
- | `reduce_dimensions.py` | PCA + variance plot, neighbors, UMAP, optional t-SNE | `python scripts/reduce_dimensions.py norm.h5ad -o red.h5ad` |
64
- | `batch_correct.py` | Integration: harmony / bbknn / combat | `python scripts/batch_correct.py red.h5ad -o int.h5ad --method harmony --batch-key sample` |
65
- | `cluster.py` | Leiden (or louvain) at one or many resolutions | `python scripts/cluster.py red.h5ad -o clu.h5ad --resolution 0.3 0.6 1.0` |
66
- | `find_markers.py` | `rank_genes_groups` + per-group CSVs + marker plots | `python scripts/find_markers.py clu.h5ad --groupby leiden -o clu.h5ad` |
67
- | `annotate.py` | Map clusters → cell types from JSON/CSV; optional marker reference dotplot | `python scripts/annotate.py clu.h5ad -o ann.h5ad --mapping map.json` |
68
- | `score_genes.py` | Score gene signatures (JSON) and/or cell-cycle phase | `python scripts/score_genes.py ann.h5ad -o scored.h5ad --gene-sets sigs.json` |
69
- | `pseudobulk.py` | Aggregate counts by sample × cell type → matrix for pydeseq2 | `python scripts/pseudobulk.py ann.h5ad --by sample cell_type --out-prefix pb` |
70
- | `subset.py` | Subset by obs values or gene list (optionally clear stale embeddings) | `python scripts/subset.py ann.h5ad -o tcells.h5ad --obs cell_type --keep "T cells"` |
71
- | `plot.py` | Generate umap/tsne/pca/violin/dotplot/heatmap/etc. from a processed object | `python scripts/plot.py ann.h5ad --kind dotplot --genes CD3D CD14 --groupby cell_type` |
72
-
73
- ### One-shot end-to-end run
74
-
75
- ```bash
76
- # Counts → clustered, marker-annotated object + figures + marker CSVs
77
- python scripts/run_pipeline.py raw.h5ad -o processed.h5ad \
78
- --resolution 0.5 --n-top-genes 2000 --scrublet
79
- # With multi-sample integration:
80
- python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --batch-key sample --batch-method harmony
81
- # Reproducible parameters via JSON (keys mirror flag names with underscores):
82
- python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --config params.json
83
- ```
84
-
85
- ### Step-by-step chain (when you need to inspect/iterate between stages)
86
-
87
- ```bash
88
- python scripts/qc_analysis.py raw.h5ad -o qc.h5ad --scrublet
89
- python scripts/preprocess.py qc.h5ad -o norm.h5ad --n-top-genes 2000
90
- python scripts/reduce_dimensions.py norm.h5ad -o red.h5ad --n-pcs 40
91
- python scripts/cluster.py red.h5ad -o clu.h5ad --resolution 0.3 0.5 0.8
92
- python scripts/find_markers.py clu.h5ad -o clu.h5ad --groupby leiden --use-raw
93
- # inspect results/markers/*.csv, decide labels, write a mapping JSON, then:
94
- python scripts/annotate.py clu.h5ad -o ann.h5ad --mapping celltypes.json
95
- ```
96
-
97
- The sections below document the underlying scanpy calls each script performs — read them when customizing beyond the script flags.
98
-
99
- ## Quick Start
100
-
101
- ### Basic Import and Setup
102
-
103
- ```python
104
- import scanpy as sc
105
- import pandas as pd
106
- import numpy as np
107
-
108
- # Configure settings
109
- sc.settings.verbosity = 3
110
- sc.settings.set_figure_params(dpi=80, facecolor='white')
111
- sc.settings.figdir = './figures/'
112
- sc.settings.autosave = True # Preferred over per-plot save= (deprecated in scanpy 1.12)
113
- ```
114
-
115
- ### Loading Data
116
-
117
- ```python
118
- # From 10X Genomics
119
- adata = sc.read_10x_mtx('path/to/data/')
120
- adata = sc.read_10x_h5('path/to/data.h5')
121
-
122
- # From h5ad (AnnData format)
123
- adata = sc.read_h5ad('path/to/data.h5ad')
124
-
125
- # From CSV
126
- adata = sc.read_csv('path/to/data.csv')
127
- ```
128
-
129
- For R-native files, do not try to parse Seurat `.rds` directly in Python. Convert first:
130
-
131
- ```bash
132
- # See references/r_interop.md for installing R and conversion packages.
133
- Rscript convert_rds_to_h5ad.R input.rds output.h5ad
134
- ```
135
-
136
- ```python
137
- adata = sc.read_h5ad('output.h5ad')
138
- ```
139
-
140
- ### Understanding AnnData Structure
141
-
142
- The AnnData object is the core data structure in scanpy:
143
-
144
- ```python
145
- adata.X # Expression matrix (cells × genes)
146
- adata.obs # Cell metadata (DataFrame)
147
- adata.var # Gene metadata (DataFrame)
148
- adata.uns # Unstructured annotations (dict)
149
- adata.obsm # Multi-dimensional cell data (PCA, UMAP)
150
- adata.raw # Raw data backup
151
-
152
- # Access cell and gene names
153
- adata.obs_names # Cell barcodes
154
- adata.var_names # Gene names
155
- ```
156
-
157
- ## Standard Analysis Workflow
158
-
159
- The seven steps, with code and the parameters that matter at each, are in
160
- [references/analysis_workflow.md](references/analysis_workflow.md):
161
-
162
- 1. **Quality control** — filter cells and genes; inspect mitochondrial fraction and counts
163
- before choosing thresholds rather than copying defaults.
164
- 2. **Normalization and preprocessing** — normalize, log-transform, select highly variable
165
- genes, and keep `.raw` for later plotting.
166
- 3. **Dimensionality reduction** — PCA, then the neighbour graph, then UMAP.
167
- 4. **Clustering** — Leiden at a resolution chosen for the question, not the default.
168
- 5. **Marker gene identification** — ranked genes per cluster.
169
- 6. **Cell type annotation** — mapping clusters to types from markers.
170
- 7. **Save results** — writing the annotated `AnnData`.
171
-
172
- Common follow-on tasks — publication plots, trajectory inference, pseudobulk differential
173
- expression between conditions, gene set scoring, and batch correction — are in the same
174
- file. See also [references/standard_workflow.md](references/standard_workflow.md) and
175
- [references/plotting_guide.md](references/plotting_guide.md).
176
-
177
- ## Key Parameters to Adjust
178
-
179
- ### Quality Control
180
- - `min_genes`: Minimum genes per cell (typically 200-500)
181
- - `min_cells`: Minimum cells per gene (typically 3-10)
182
- - `pct_counts_mt`: Mitochondrial threshold (typically 5-20%)
183
-
184
- ### Normalization
185
- - `target_sum`: Target counts per cell (default 1e4)
186
-
187
- ### Feature Selection
188
- - `n_top_genes`: Number of HVGs (typically 2000-3000)
189
- - `min_mean`, `max_mean`, `min_disp`: HVG selection parameters
190
-
191
- ### Dimensionality Reduction
192
- - `n_pcs`: Number of principal components (check variance ratio plot)
193
- - `n_neighbors`: Number of neighbors (typically 10-30)
194
-
195
- ### Clustering
196
- - `resolution`: Clustering granularity (0.4-1.2, higher = more clusters)
197
-
198
- ## Common Pitfalls and Best Practices
199
-
200
- 1. **Always save raw counts**: `adata.raw = adata` before filtering genes
201
- 2. **Check QC plots carefully**: Adjust thresholds based on dataset quality
202
- 3. **Use Leiden clustering**: `sc.tl.louvain` is deprecated in scanpy 1.12
203
- 4. **Try multiple clustering resolutions**: Find optimal granularity
204
- 5. **Validate cell type annotations**: Use multiple marker genes
205
- 6. **Use `use_raw=True` for gene expression plots**: Shows normalized counts from `.raw`
206
- 7. **Check PCA variance ratio**: Determine optimal number of PCs
207
- 8. **Save intermediate results**: Long workflows can fail partway through
208
- 9. **Pseudobulk for DE**: Do not treat `rank_genes_groups` p-values as rigorous DE between conditions
209
- 10. **Save plots via settings**: Use `sc.settings.autosave` instead of deprecated `save=` on plot functions
210
- 11. **Convert R objects before Scanpy**: Use R packages to convert Seurat or SingleCellExperiment `.rds` files to `.h5ad`, preserving counts, metadata, and gene identifiers
211
-
212
- ## Bundled Resources
213
-
214
- ### scripts/ (CLI toolkit)
215
- A composable set of `.h5ad`-in/`.h5ad`-out scripts covering the whole workflow plus a one-command end-to-end pipeline. See the **Script Toolkit** section above for the full table and chaining examples. Each script has `--help`. Files:
216
-
217
- - `_common.py` — shared loading/saving/figure helpers imported by the others (not a CLI)
218
- - `run_pipeline.py` — full pipeline in one command (flags or `--config` JSON)
219
- - `inspect_data.py`, `convert.py` — explore and load/convert any input format
220
- - `qc_analysis.py`, `preprocess.py`, `reduce_dimensions.py`, `batch_correct.py`, `cluster.py` — pipeline steps
221
- - `find_markers.py`, `annotate.py`, `score_genes.py`, `pseudobulk.py` — markers, annotation, scoring, DE prep
222
- - `subset.py`, `plot.py` — subset by metadata/genes; generate any standard plot
223
-
224
- **Default to these scripts before writing scanpy code from scratch.**
225
-
226
- ### references/standard_workflow.md
227
- Complete step-by-step workflow with detailed explanations and code examples for:
228
- - Data loading and setup
229
- - Quality control with visualization
230
- - Normalization and scaling
231
- - Feature selection
232
- - Dimensionality reduction (PCA, UMAP, t-SNE)
233
- - Clustering (Leiden)
234
- - Doublet detection (scrublet) and pseudobulk aggregation
235
- - Marker gene identification
236
- - Cell type annotation
237
- - Trajectory inference
238
- - Differential expression
239
-
240
- Read this reference when performing a complete analysis from scratch.
241
-
242
- ### references/api_reference.md
243
- Quick reference guide for scanpy functions organized by module:
244
- - Reading/writing data (`sc.read_*`, `adata.write_*`)
245
- - Preprocessing (`sc.pp.*`)
246
- - Tools (`sc.tl.*`)
247
- - Plotting (`sc.pl.*`)
248
- - AnnData structure and manipulation
249
- - Settings and utilities
250
-
251
- Use this for quick lookup of function signatures and common parameters.
252
-
253
- ### references/plotting_guide.md
254
- Comprehensive visualization guide including:
255
- - Quality control plots
256
- - Dimensionality reduction visualizations
257
- - Clustering visualizations
258
- - Marker gene plots (heatmaps, dot plots, violin plots)
259
- - Trajectory and pseudotime plots
260
- - Publication-quality customization
261
- - Multi-panel figures
262
- - Color palettes and styling
263
-
264
- Consult this when creating publication-ready figures.
265
-
266
- ### references/r_interop.md
267
- Agent runbook for installing R on macOS, Linux, and Windows, installing CRAN/Bioconductor conversion packages, inspecting `.rds`/`.RData` inputs, converting Seurat or SingleCellExperiment objects to `.h5ad`, and validating the result in Scanpy.
268
-
269
- ### assets/analysis_template.py
270
- Complete analysis template providing a full workflow from data loading through cell type annotation. Copy and customize this template for new analyses:
271
-
272
- ```bash
273
- cp assets/analysis_template.py my_analysis.py
274
- # Edit parameters and run
275
- python my_analysis.py
276
- ```
277
-
278
- The template includes all standard steps with configurable parameters and helpful comments.
279
-
280
- ### assets/ JSON templates
281
- Edit-and-pass templates so you don't author config/mappings from scratch:
282
- - `assets/pipeline_config.json` — parameter set for `run_pipeline.py --config`
283
- - `assets/celltype_mapping.json` — cluster → cell-type map for `annotate.py --mapping`
284
- - `assets/gene_signatures.json` — gene-set signatures for `score_genes.py --gene-sets`
285
-
286
- ## Additional Resources
287
-
288
- - **Official scanpy documentation**: https://scanpy.scverse.org/en/stable/
289
- - **Scanpy tutorials**: https://scanpy.scverse.org/en/stable/tutorials/index.html
290
- - **Release notes**: https://scanpy.scverse.org/en/stable/release-notes/index.html
291
- - **scverse ecosystem**: https://scverse.org/ (related tools: squidpy, scvi-tools, cellrank)
292
- - **R interoperability**: https://www.bioconductor.org/packages/release/bioc/html/zellkonverter.html and https://mojaveazure.github.io/seurat-disk/
293
- - **Best practices**: Luecken & Theis (2019) "Current best practices in single-cell RNA-seq"
294
-
295
- ## Tips for Effective Analysis
296
-
297
- 1. **Start with the template**: Use `assets/analysis_template.py` as a starting point
298
- 2. **Run QC script first**: Use `scripts/qc_analysis.py` for initial filtering
299
- 3. **Consult references as needed**: Load workflow and API references into context
300
- 4. **Iterate on clustering**: Try multiple resolutions and visualization methods
301
- 5. **Validate biologically**: Check marker genes match expected cell types
302
- 6. **Document parameters**: Record QC thresholds and analysis settings
303
- 7. **Save checkpoints**: Write intermediate results at key steps
@@ -1,296 +0,0 @@
1
- ---
2
- name: scholar-evaluation
3
- description: Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls. Never use for ranking people or consequential decisions.
4
- license: MIT
5
- compatibility: Requires Python 3.11+ for optional bundled standard-library CLIs. All tooling is local JSON/CSV processing with no network, credentials, external models, or subprocesses.
6
- allowed-tools: Read Write Bash Glob Python
7
- metadata:
8
- version: "2.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # Scholar Evaluation
13
-
14
- ## Purpose
15
-
16
- Provide developmental, evidence-traceable feedback on a **scholarly work**:
17
- paper, draft, protocol, literature synthesis, or research idea. Use
18
- qualitative judgment first. Optional scores only describe how submitted
19
- evidence maps to a predeclared bounded rubric.
20
-
21
- This skill also audits whether a low-stakes assessment process documents its
22
- construct, provenance, rater quality, uncertainty, traceability, sensitivity,
23
- fairness, accessibility, privacy, and human governance.
24
-
25
- ## Hard safety boundary
26
-
27
- Never use this skill to automate, recommend, materially influence, or score:
28
-
29
- - hiring, promotion, or tenure;
30
- - admissions;
31
- - grants or other funding;
32
- - prizes, honors, or awards;
33
- - discipline, dismissal, or sanctions; or
34
- - any other high-impact personnel decision.
35
-
36
- Never rank people. Never reduce a person to a composite score. Never infer
37
- ability, character, integrity, protected traits, future performance, or worth.
38
- A nominal human-in-the-loop does not remove this boundary.
39
-
40
- If asked for a prohibited use, stop. Offer developmental comments on a
41
- scholarly work or a process-only audit that does not process applications,
42
- compare people, recommend an outcome, or advise a decision.
43
-
44
- Do not issue publication-readiness, accept/reject, or “top-tier” judgments.
45
-
46
- Read `references/responsible_assessment.md` before any organizational use.
47
-
48
- ## ScholarEval status
49
-
50
- The referenced ScholarEval project is an **experimental
51
- literature-grounded research-idea evaluation framework**, not validated
52
- psychometrics.
53
-
54
- The verified primary record is Moussa et al., *ScholarEval: Research Idea
55
- Evaluation Grounded in Literature*, arXiv:2510.16234v2, revised 2026-02-28.
56
- It reports a retrieval-augmented soundness/contribution framework, a
57
- 117-idea four-discipline dataset, coverage experiments, and a user study.
58
-
59
- Do not generalize those results to person assessment, consequential decisions,
60
- all disciplines, or this skill's rubric. No peer-reviewed publication status
61
- was verified during the dated review. See `references/source_ledger.md`.
62
-
63
- ## Metric and prestige policy
64
-
65
- Do not score or infer quality from:
66
-
67
- - Journal Impact Factor or other journal measures;
68
- - h-index, publication counts, or citation counts;
69
- - altmetrics or attention;
70
- - journal, conference, venue, institution, employer, or geographic prestige;
71
- - author affiliation, reputation, network, or career path.
72
-
73
- The rubric validator rejects common proxy-measure criteria.
74
-
75
- If a qualified reviewer mentions an indicator descriptively outside the
76
- scoring tools, record its exact purpose, source, coverage, field and time
77
- effects, uncertainty, missingness, biases, gaming risk, and why it does not
78
- directly measure quality. Never hide indicators inside an opaque composite.
79
-
80
- ## Data boundary
81
-
82
- Bundled scripts accept only strict local JSON/CSV containing pseudonymous IDs,
83
- bounded ratings, statuses, uncertainty, and local references.
84
-
85
- Do not put raw private applications, CVs, letters, reviewer identities,
86
- contact details, protected attributes, or source-document text in inputs,
87
- outputs, logs, examples, or prompts. Keep source content in the authorized
88
- records system and use opaque local references.
89
-
90
- Allowed classifications are:
91
-
92
- - `synthetic`
93
- - `public_scholarly_work`
94
- - `deidentified_low_stakes`
95
-
96
- No script searches the web, loads environment files, reads credentials, calls a
97
- model, executes supplied text, deserializes executable objects, or launches a
98
- process.
99
-
100
- Use Bash only to invoke the documented local `python3` commands.
101
-
102
- ## Workflow
103
-
104
- ### 1. Confirm allowed use and authorization
105
-
106
- Record:
107
-
108
- - developmental purpose;
109
- - unit of assessment: `scholarly_work`;
110
- - work type, stage, discipline, language, and audience;
111
- - authorized source location and data classification;
112
- - accountable committee owner;
113
- - conflicts and recusals;
114
- - accessibility and accommodation process;
115
- - appeal or correction route; and
116
- - data purpose, access, retention, and deletion.
117
-
118
- Stop on a prohibited decision context or unnecessary private data.
119
-
120
- ### 2. Define the construct before criteria
121
-
122
- State:
123
-
124
- - what quality or support is being examined;
125
- - excluded constructs;
126
- - intended interpretation;
127
- - contexts where the interpretation does not travel;
128
- - evidence requirements; and
129
- - known limitations.
130
-
131
- Start with values and disciplinary context, not available metrics.
132
-
133
- ### 3. Adapt and validate the rubric
134
-
135
- Begin with `assets/rubric_template.json`, then obtain qualified disciplinary,
136
- assessment-methods, stakeholder, accessibility, privacy, and fairness review.
137
-
138
- The template deliberately records content validity as `not_established`.
139
- Do not change that status without documented evidence for the exact intended
140
- use.
141
-
142
- Validate structure:
143
-
144
- ```bash
145
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/validate_rubric.py \
146
- --rubric assets/rubric_template.json
147
- ```
148
-
149
- Read `references/evaluation_framework.md` for construct, anchor, validity, and
150
- rater guidance.
151
-
152
- ### 4. Build traceable evidence records
153
-
154
- Reviewers may read an authorized work outside the scripts. Record only stable
155
- local locators and claim references in
156
- `assets/evidence_manifest_template.json`.
157
-
158
- For every criterion, distinguish:
159
-
160
- - observed evidence from interpretation;
161
- - supporting from contrary evidence;
162
- - available from unavailable evidence;
163
- - `missing` from `not_applicable`; and
164
- - uncertainty from absence.
165
-
166
- Failure to find prior work does not prove novelty.
167
-
168
- ### 5. Rate independently
169
-
170
- Use `assets/evaluation_template.json`. Each criterion must be:
171
-
172
- - `rated` with an anchor score, bounded uncertainty, evidence IDs, and a local
173
- rationale reference;
174
- - `missing` with null score/uncertainty and a rationale reference; or
175
- - `not_applicable` with null score/uncertainty and a rationale reference.
176
-
177
- Do not encode missing or not-applicable as zero. Raters should train, calibrate,
178
- disclose conflicts, rate independently, and document disagreement.
179
-
180
- ### 6. Run local quality checks
181
-
182
- Bounded scoring, without labels or recommendation:
183
-
184
- ```bash
185
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/calculate_scores.py \
186
- --rubric assets/rubric_template.json \
187
- --evaluation assets/evaluation_template.json
188
- ```
189
-
190
- Evidence traceability:
191
-
192
- ```bash
193
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_traceability.py \
194
- --rubric assets/rubric_template.json \
195
- --evaluation assets/evaluation_template.json \
196
- --evidence assets/evidence_manifest_template.json
197
- ```
198
-
199
- Inter-rater agreement:
200
-
201
- ```bash
202
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/summarize_agreement.py \
203
- --rubric assets/rubric_template.json \
204
- --ratings assets/ratings_template.csv
205
- ```
206
-
207
- Weight sensitivity requires two or more distinct scholarly-work evaluation
208
- files:
209
-
210
- ```bash
211
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/weight_sensitivity.py \
212
- --rubric assets/rubric_template.json \
213
- --evaluation /tmp/work-a-evaluation.json \
214
- --evaluation /tmp/work-b-evaluation.json
215
- ```
216
-
217
- Process controls:
218
-
219
- ```bash
220
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/check_process.py \
221
- --process assets/process_checklist_template.json
222
- ```
223
-
224
- The checklist template is intentionally unconfirmed and fails closed.
225
- Instructions and exact schemas are in `references/local_tooling.md`.
226
-
227
- ### 7. Synthesize qualitative findings
228
-
229
- Lead with criterion-level evidence, not the composite. For each criterion:
230
-
231
- 1. cite evidence references;
232
- 2. state `rated`, `missing`, or `not_applicable`;
233
- 3. explain the anchor interpretation;
234
- 4. report score and uncertainty only if rated;
235
- 5. note disagreements and context;
236
- 6. identify strengths and limitations; and
237
- 7. offer non-prescriptive improvement options.
238
-
239
- Generate an empty-reference scaffold if useful:
240
-
241
- ```bash
242
- PYTHONDONTWRITEBYTECODE=1 python3 scripts/generate_report_scaffold.py \
243
- --rubric assets/rubric_template.json \
244
- --evaluation assets/evaluation_template.json \
245
- --output /tmp/developmental-report-scaffold.json
246
- ```
247
-
248
- The scaffold does not read source documents or draft findings.
249
-
250
- ### 8. Human review and release
251
-
252
- Before releasing an organizational report, a qualified accountable human
253
- committee must verify:
254
-
255
- - construct and rubric provenance;
256
- - content-validity evidence and limits;
257
- - rater training, agreement, inter-rater reliability evidence, and drift;
258
- - evidence traceability and source access;
259
- - missingness, not-applicable rationales, and uncertainty;
260
- - weight sensitivity and order instability;
261
- - disciplinary and subgroup bias review;
262
- - conflicts and recusals;
263
- - accessibility and accommodations;
264
- - privacy, minimization, retention, and output controls; and
265
- - correction or appeal information.
266
-
267
- Document dissent. Do not imply consensus, validity, or precision beyond the
268
- evidence. Periodically evaluate the evaluation and retire harmful criteria.
269
-
270
- ## Interpretation rules
271
-
272
- - A score is an ordinal rubric summary, not a natural measurement.
273
- - Normalization does not repair incomplete evidence.
274
- - The bundled uncertainty range is not a confidence interval.
275
- - Agreement does not establish reliability, validity, fairness, or correctness.
276
- - Stable results under tested weights do not establish validity.
277
- - The overall score never overrides criterion evidence or qualified judgment.
278
- - No output is a decision recommendation.
279
-
280
- ## Bundled resources
281
-
282
- - `references/responsible_assessment.md` — safety, metrics, governance,
283
- accessibility, privacy, and bias.
284
- - `references/evaluation_framework.md` — ScholarEval boundary, construct,
285
- criteria, anchors, validity, and interpretation.
286
- - `references/local_tooling.md` — strict schemas, formulas, commands, and
287
- output behavior.
288
- - `references/source_ledger.md` — authoritative sources and publication-status
289
- verification dated 2026-07-23.
290
- - `references/security_validation.md` — baseline remediation, validation, and
291
- residual security-scan record.
292
- - `assets/rubric_template.json` — bounded rubric template.
293
- - `assets/evaluation_template.json` — rating template.
294
- - `assets/evidence_manifest_template.json` — traceability template.
295
- - `assets/process_checklist_template.json` — fail-closed process checklist.
296
- - `assets/ratings_template.csv` — synthetic agreement data.