@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,412 +0,0 @@
1
- ---
2
- name: neuropixels-analysis
3
- description: Analyze Neuropixels extracellular recordings end-to-end with SpikeInterface. Covers loading SpikeGLX/Open Ephys/NWB data, preprocessing, drift/motion correction, Kilosort4 (and CPU) spike sorting, quality metrics, and unit curation (threshold-based, model-based UnitRefine, and AI-assisted visual review). Use when working with Neuropixels 1.0/2.0 recordings, spike sorting, or extracellular electrophysiology analysis.
4
- license: MIT license
5
- metadata:
6
- version: "2.3"
7
- skill-author: K-Dense Inc.
8
- openclaw:
9
- primaryEnv: ANTHROPIC_API_KEY
10
- envVars:
11
- - name: ANTHROPIC_API_KEY
12
- required: false
13
- description: For optional Claude API calls.
14
- ---
15
-
16
- # Neuropixels Data Analysis
17
-
18
- ## Overview
19
-
20
- Toolkit for analyzing Neuropixels high-density neural recordings using current best
21
- practices from [SpikeInterface](https://spikeinterface.readthedocs.io/), the Allen
22
- Institute, and the International Brain Laboratory (IBL). It covers the full workflow from
23
- raw data to publication-ready curated units.
24
-
25
- All examples use the real SpikeInterface API (`spikeinterface.full as si`) plus the
26
- companion curation module (`spikeinterface.curation as sc`). The skill ships runnable
27
- scripts in `scripts/` and a copy-and-edit template in `assets/` that implement this
28
- workflow directly on top of SpikeInterface — there is no separate package to install
29
- beyond the dependencies listed under [Installation](#installation).
30
-
31
- ## When to Use This Skill
32
-
33
- This skill should be used when:
34
- - Working with Neuropixels recordings (`.ap.bin`, `.lf.bin`, `.meta` files)
35
- - Loading data from SpikeGLX, Open Ephys, or NWB formats
36
- - Preprocessing neural recordings (filtering, common reference, bad-channel detection)
37
- - Detecting and correcting motion/drift
38
- - Running spike sorting (Kilosort4, SpykingCircus2, Mountainsort5, Tridesclous2)
39
- - Computing quality metrics (SNR, ISI violations, presence ratio, amplitude cutoff)
40
- - Curating units (threshold-based, model-based, or AI-assisted)
41
- - Creating visualizations and exporting to Phy or NWB
42
-
43
- ## Supported Hardware & Formats
44
-
45
- | Probe | Electrodes | Channels | Notes |
46
- |-------|-----------|----------|-------|
47
- | Neuropixels 1.0 | 960 | 384 | Use `phase_shift` for ADC correction |
48
- | Neuropixels 2.0 (single) | 1280 | 384 | Denser geometry |
49
- | Neuropixels 2.0 (4-shank) | 5120 | 384 | Multi-region recording |
50
-
51
- | Format | Extension | Reader |
52
- |--------|-----------|--------|
53
- | SpikeGLX | `.ap.bin`, `.lf.bin`, `.meta` | `si.read_spikeglx()` |
54
- | Open Ephys | `.continuous`, `.oebin` | `si.read_openephys()` |
55
- | NWB | `.nwb` | `si.read_nwb()` |
56
-
57
- ## Quick Start
58
-
59
- ### Import and configure parallel processing
60
-
61
- ```python
62
- import spikeinterface.full as si
63
-
64
- # Global job kwargs are reused by all parallelizable steps
65
- si.set_global_job_kwargs(n_jobs=-1, chunk_duration="1s", progress_bar=True)
66
- ```
67
-
68
- ### Loading data
69
-
70
- ```python
71
- # Inspect available streams first
72
- stream_names, stream_ids = si.get_neo_streams("spikeglx", "/path/to/run_g0/")
73
- print(stream_names) # e.g. ['imec0.ap', 'imec0.lf', 'nidq']
74
-
75
- # SpikeGLX (most common) — select the AP stream by name
76
- recording = si.read_spikeglx("/path/to/run_g0/", stream_name="imec0.ap", load_sync_channel=False)
77
-
78
- # Open Ephys
79
- recording = si.read_openephys("/path/to/Record_Node_101/")
80
-
81
- # For quick iteration, slice the first 60 s
82
- fs = recording.get_sampling_frequency()
83
- recording_sub = recording.frame_slice(0, int(60 * fs))
84
- ```
85
-
86
- ### Full pipeline (bundled script)
87
-
88
- The repository ships an end-to-end pipeline built on SpikeInterface:
89
-
90
- ```bash
91
- python scripts/neuropixels_pipeline.py /path/to/spikeglx/data output/ --sorter kilosort4 --curation allen
92
- ```
93
-
94
- It performs load → preprocess → drift check → optional motion correction → sorting →
95
- postprocessing → quality metrics → curation → export. Read the steps below to run them
96
- interactively or customize the pipeline.
97
-
98
- ## Standard Analysis Workflow
99
-
100
- ### 1. Preprocessing
101
-
102
- Recommended chain, following the SpikeInterface Neuropixels how-to (IBL-style destriping
103
- with channel removal + common reference):
104
-
105
- ```python
106
- rec = si.highpass_filter(recording, freq_min=400.0)
107
- bad_channel_ids, channel_labels = si.detect_bad_channels(rec)
108
- rec = rec.remove_channels(bad_channel_ids)
109
- rec = si.phase_shift(rec) # ADC phase correction (Neuropixels 1.0)
110
- rec = si.common_reference(rec, operator="median", reference="global")
111
- ```
112
-
113
- Save the preprocessed recording (Kilosort needs a binary file, and it speeds up reuse):
114
-
115
- ```python
116
- rec = rec.save(folder="preprocessed/", format="binary")
117
- ```
118
-
119
- ### 2. Check and correct drift
120
-
121
- Always inspect drift before sorting:
122
-
123
- ```python
124
- from spikeinterface.sortingcomponents.peak_detection import detect_peaks
125
- from spikeinterface.sortingcomponents.peak_localization import localize_peaks
126
-
127
- noise_levels = si.get_noise_levels(rec, return_in_uV=False)
128
- peaks = detect_peaks(rec, method="locally_exclusive", noise_levels=noise_levels,
129
- detect_threshold=5, radius_um=50.0)
130
- peak_locations = localize_peaks(rec, peaks, method="center_of_mass")
131
-
132
- # Visualize the drift raster
133
- si.plot_drift_raster_map(peaks=peaks, peak_locations=peak_locations,
134
- recording=rec, clim=(-50, 50))
135
- ```
136
-
137
- Apply correction if needed (presets: `rigid_fast`, `kilosort_like`,
138
- `nonrigid_accurate`, `nonrigid_fast_and_accurate`, `dredge`, `dredge_fast`):
139
-
140
- ```python
141
- rec_corrected = si.correct_motion(rec, preset="nonrigid_fast_and_accurate", folder="motion/")
142
- ```
143
-
144
- ### 3. Spike sorting
145
-
146
- ```python
147
- # Kilosort4 (recommended, requires a CUDA GPU)
148
- sorting = si.run_sorter("kilosort4", rec_corrected, folder="ks4_output")
149
-
150
- # CPU alternatives (internally developed, no external install)
151
- sorting = si.run_sorter("spykingcircus2", rec_corrected, folder="sc2_output")
152
- sorting = si.run_sorter("tridesclous2", rec_corrected, folder="tdc2_output")
153
- sorting = si.run_sorter("mountainsort5", rec_corrected, folder="ms5_output")
154
-
155
- # External sorters can run in containers without local install
156
- sorting = si.run_sorter("kilosort2_5", rec_corrected, folder="ks25_output", docker_image=True)
157
-
158
- print(si.installed_sorters())
159
- ```
160
-
161
- > Note: `run_sorter` uses the `folder=` argument. The older `output_folder=` is deprecated.
162
-
163
- ### 4. Postprocessing
164
-
165
- ```python
166
- analyzer = si.create_sorting_analyzer(sorting, rec_corrected, sparse=True,
167
- format="binary_folder", folder="analyzer/")
168
-
169
- analyzer.compute("random_spikes", method="uniform", max_spikes_per_unit=500)
170
- analyzer.compute("waveforms", ms_before=1.0, ms_after=2.0)
171
- analyzer.compute("templates", operators=["average", "std"])
172
- analyzer.compute("noise_levels")
173
- analyzer.compute("spike_amplitudes")
174
- analyzer.compute("correlograms", window_ms=50.0, bin_ms=1.0)
175
- analyzer.compute("unit_locations", method="monopolar_triangulation")
176
- analyzer.compute("template_similarity")
177
-
178
- metric_names = ["firing_rate", "presence_ratio", "snr", "isi_violation", "amplitude_cutoff"]
179
- analyzer.compute("quality_metrics", metric_names=metric_names)
180
- metrics = analyzer.get_extension("quality_metrics").get_data()
181
- ```
182
-
183
- ### 5. Curation by metric thresholds
184
-
185
- ```python
186
- # Allen-style query (note: column is isi_violations_ratio)
187
- query = "(amplitude_cutoff < 0.1) & (isi_violations_ratio < 0.5) & (presence_ratio > 0.9)"
188
- good_unit_ids = metrics.query(query).index.values
189
- ```
190
-
191
- For reusable, multi-threshold logic with `allen` / `ibl` / `strict` presets, use the
192
- bundled `scripts/compute_metrics.py`. See
193
- [references/AUTOMATED_CURATION.md](references/AUTOMATED_CURATION.md) for details and the
194
- Bombcell / UnitMatch tools.
195
-
196
- ### 6. Model-based curation (UnitRefine)
197
-
198
- SpikeInterface can apply pretrained machine-learning classifiers from Hugging Face via the
199
- `spikeinterface.curation` module. The UnitRefine models were trained on real Neuropixels
200
- data (V1, SC, ALM):
201
-
202
- ```python
203
- import spikeinterface.curation as sc
204
-
205
- # 1) noise vs neural
206
- noise_labels = sc.model_based_label_units(
207
- sorting_analyzer=analyzer,
208
- repo_id="SpikeInterface/UnitRefine_noise_neural_classifier",
209
- trust_model=True,
210
- )
211
- neural = analyzer.remove_units(noise_labels[noise_labels["prediction"] == "noise"].index)
212
-
213
- # 2) single-unit (sua) vs multi-unit (mua) on the surviving units
214
- sua_mua_labels = sc.model_based_label_units(
215
- sorting_analyzer=neural,
216
- repo_id="SpikeInterface/UnitRefine_sua_mua_classifier",
217
- trust_model=True,
218
- )
219
- ```
220
-
221
- Each call returns a DataFrame with `prediction` and `probability` (confidence) per unit.
222
- `trust_model=True` (or an explicit `trusted=[...]` list) is required to load the `.skops`
223
- model — only load models from sources you trust. Models trained on other brain
224
- areas/datasets may not transfer; validate against a manually labelled subset.
225
-
226
- ### 7. AI-assisted curation (for uncertain units)
227
-
228
- When running inside an agent such as Cursor or Claude Code, the agent can directly inspect
229
- waveform/correlogram plots and give an expert read — no API setup required. Generate plots
230
- and ask the agent to assess isolation quality.
231
-
232
- For programmatic vision-model access, **read API keys from the environment — never hardcode
233
- credentials in analysis scripts** (they leak into version control and logs):
234
-
235
- ```python
236
- import os
237
- from anthropic import Anthropic
238
-
239
- client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"]) # set this in your shell, not in code
240
- ```
241
-
242
- See [references/AI_CURATION.md](references/AI_CURATION.md) for the full pattern (rendering a
243
- unit summary image, building the prompt, and parsing the response).
244
-
245
- ### 8. Export results
246
-
247
- ```python
248
- # Keep only good units, then export
249
- analyzer_clean = analyzer.select_units(good_unit_ids, folder="analyzer_clean/", format="binary_folder")
250
-
251
- # Phy for manual review
252
- si.export_to_phy(analyzer_clean, output_folder="phy_export/",
253
- compute_pc_features=True, compute_amplitudes=True)
254
-
255
- # Figures report
256
- si.export_report(analyzer_clean, "report/", format="png")
257
-
258
- # NWB
259
- from spikeinterface.exporters import export_to_nwb
260
- export_to_nwb(analyzer_clean, "output.nwb")
261
-
262
- # Metrics table
263
- metrics.to_csv("quality_metrics.csv")
264
- ```
265
-
266
- ## Common Pitfalls and Best Practices
267
-
268
- 1. **Always check drift** before spike sorting — drift > ~10 μm meaningfully degrades quality.
269
- 2. **Use `phase_shift`** for Neuropixels 1.0 to correct ADC sampling offsets.
270
- 3. **Save the preprocessed recording** with `rec.save(folder=...)` to avoid recomputation (Kilosort also needs a binary file).
271
- 4. **Use a GPU** for Kilosort4 — it is far faster than CPU sorters.
272
- 5. **Review uncertain units** — automated/model-based curation is a starting point, not a verdict.
273
- 6. **Combine approaches** — thresholds for clear cases, model/AI for borderline units.
274
- 7. **Document thresholds and model repo IDs** for reproducibility.
275
- 8. **Export to Phy** for critical experiments — human oversight is valuable.
276
-
277
- ## Key Parameters to Adjust
278
-
279
- ### Preprocessing
280
- - `freq_min`: highpass cutoff (300–400 Hz typical)
281
- - `detect_bad_channels`: returns `(bad_channel_ids, channel_labels)`
282
-
283
- ### Motion Correction
284
- - `preset`: `nonrigid_fast_and_accurate` (balanced), `nonrigid_accurate` (severe drift), `dredge` (state of the art)
285
-
286
- ### Spike Sorting (Kilosort4)
287
- - `batch_size`: samples per batch (60000 default)
288
- - `nblocks`: drift blocks (increase for long, drifty recordings)
289
- - `Th_universal` / `Th_learned`: detection thresholds (lower = more spikes)
290
-
291
- ### Quality Metrics
292
- - `snr`: signal-to-noise cutoff (3–5 typical)
293
- - `isi_violations_ratio`: refractory violations (0.01–0.5)
294
- - `presence_ratio`: recording coverage (0.5–0.95)
295
-
296
- ## Bundled Resources
297
-
298
- ### scripts/explore_recording.py
299
- Quick inspection of a recording (streams, channels, duration, bad channels):
300
- ```bash
301
- python scripts/explore_recording.py /path/to/data
302
- ```
303
-
304
- ### scripts/preprocess_recording.py
305
- Automated preprocessing:
306
- ```bash
307
- python scripts/preprocess_recording.py /path/to/data --output preprocessed/
308
- ```
309
-
310
- ### scripts/run_sorting.py
311
- Run spike sorting:
312
- ```bash
313
- python scripts/run_sorting.py preprocessed/ --sorter kilosort4 --output sorting/
314
- ```
315
-
316
- ### scripts/compute_metrics.py
317
- Compute quality metrics and apply curation:
318
- ```bash
319
- python scripts/compute_metrics.py sorting/ preprocessed/ --output metrics/ --curation allen
320
- ```
321
-
322
- ### scripts/export_to_phy.py
323
- Export to Phy for manual curation:
324
- ```bash
325
- python scripts/export_to_phy.py metrics/analyzer --output phy_export/
326
- ```
327
-
328
- ### scripts/neuropixels_pipeline.py
329
- Complete end-to-end pipeline (see [Quick Start](#full-pipeline-bundled-script)).
330
-
331
- ### assets/analysis_template.py
332
- Complete, editable analysis template. Copy and customize:
333
- ```bash
334
- cp assets/analysis_template.py my_analysis.py
335
- # Edit the PARAMETERS section, then run
336
- python my_analysis.py
337
- ```
338
-
339
- ## Detailed Reference Guides
340
-
341
- | Topic | Reference |
342
- |-------|-----------|
343
- | Full workflow | [references/standard_workflow.md](references/standard_workflow.md) |
344
- | API reference (SpikeInterface) | [references/api_reference.md](references/api_reference.md) |
345
- | Plotting guide | [references/plotting_guide.md](references/plotting_guide.md) |
346
- | Preprocessing | [references/PREPROCESSING.md](references/PREPROCESSING.md) |
347
- | Spike sorting | [references/SPIKE_SORTING.md](references/SPIKE_SORTING.md) |
348
- | Motion correction | [references/MOTION_CORRECTION.md](references/MOTION_CORRECTION.md) |
349
- | Quality metrics | [references/QUALITY_METRICS.md](references/QUALITY_METRICS.md) |
350
- | Automated & model-based curation | [references/AUTOMATED_CURATION.md](references/AUTOMATED_CURATION.md) |
351
- | AI-assisted curation | [references/AI_CURATION.md](references/AI_CURATION.md) |
352
- | Waveform analysis | [references/ANALYSIS.md](references/ANALYSIS.md) |
353
-
354
- ## Installation
355
-
356
- Requires Python ≥ 3.10. Using [uv](https://docs.astral.sh/uv/) is recommended.
357
-
358
- ```bash
359
- # Core packages (SpikeInterface bundles the curation/model tooling)
360
- uv pip install "spikeinterface[full]" probeinterface neo
361
-
362
- # Spike sorters
363
- uv pip install kilosort # Kilosort4 (CUDA GPU required)
364
- uv pip install spykingcircus # SpykingCircus (legacy; SpykingCircus2 ships with SpikeInterface)
365
- uv pip install mountainsort5 # Mountainsort5 (CPU)
366
-
367
- # Model-based curation (UnitRefine) downloads from Hugging Face
368
- uv pip install "huggingface_hub" skops
369
-
370
- # Optional: AI-assisted visual curation
371
- uv pip install anthropic
372
-
373
- # Optional: IBL tools and Bombcell
374
- uv pip install ibl-neuropixel ibllib bombcell
375
- ```
376
-
377
- For reproducible environments, pin versions (current as of 2026-06: `spikeinterface==0.104.3`,
378
- `kilosort==4.1.7`, `probeinterface==0.3.2`, `neo==0.14.4`). Unpinned installs are fine for
379
- quick experimentation but should be pinned in production pipelines.
380
-
381
- ## Project Structure
382
-
383
- ```
384
- project/
385
- ├── raw_data/
386
- │ └── recording_g0/
387
- │ └── recording_g0_imec0/
388
- │ ├── recording_g0_t0.imec0.ap.bin
389
- │ └── recording_g0_t0.imec0.ap.meta
390
- ├── preprocessed/ # Saved preprocessed recording
391
- ├── motion/ # Motion estimation results
392
- ├── sorting_output/ # Spike sorter output
393
- ├── analyzer/ # SortingAnalyzer (waveforms, metrics)
394
- ├── phy_export/ # For manual curation
395
- ├── ai_curation/ # AI analysis reports
396
- └── results/
397
- ├── quality_metrics.csv
398
- ├── curation_labels.json
399
- └── output.nwb
400
- ```
401
-
402
- ## Additional Resources
403
-
404
- - **SpikeInterface Docs**: https://spikeinterface.readthedocs.io/
405
- - **Neuropixels Tutorial**: https://spikeinterface.readthedocs.io/en/stable/how_to/analyze_neuropixels.html
406
- - **Model-based Curation Tutorial**: https://spikeinterface.readthedocs.io/en/stable/tutorials/curation/plot_1_automated_curation.html
407
- - **UnitRefine Models (Hugging Face)**: https://huggingface.co/SpikeInterface
408
- - **Kilosort4 GitHub**: https://github.com/MouseLand/Kilosort
409
- - **IBL Neuropixel Tools**: https://github.com/int-brain-lab/ibl-neuropixel
410
- - **Allen Institute ecephys**: https://github.com/AllenInstitute/ecephys_spike_sorting
411
- - **Bombcell (Automated QC)**: https://github.com/Julie-Fabre/bombcell
412
- - **Awesome Neuropixels**: https://github.com/Julie-Fabre/awesome_neuropixels
@@ -1,195 +0,0 @@
1
- ---
2
- name: nextflow
3
- description: Build, run, and debug Nextflow data pipelines and nf-core workflows end to end. Use whenever the user mentions Nextflow, nf-core, .nf files, nextflow.config, DSL2, processes/channels/operators, samplesheets, or wants to run a community pipeline (e.g. nf-core/rnaseq, nf-core/sarek), write or test a module/subworkflow with nf-test, configure executors/containers (Docker, Singularity/Apptainer, Conda, Wave), scale a workflow to HPC/SLURM or cloud (AWS Batch, Google Batch, Azure, Kubernetes), or debug a failed/-resume run. Make sure to use this skill for any reproducible scientific/bioinformatics workflow work even if the user does not say the word "Nextflow", and for authoring nf-core-compliant pipelines, modules, configs, and linting.
4
- license: Apache-2.0
5
- metadata:
6
- version: "1.1"
7
- skill-author: K-Dense Inc.
8
- ---
9
-
10
- # Nextflow
11
-
12
- ## Overview
13
-
14
- Nextflow is a workflow language and runtime for building **reproducible, portable, scalable** data pipelines. It is dominant in bioinformatics but works for any data-heavy computation. nf-core is a community curating production-grade Nextflow pipelines, reusable modules, and the `nf-core` tooling on top of Nextflow.
15
-
16
- Key ideas:
17
- - **Dataflow programming**: pipelines are `process` tasks connected by **channels**. Nextflow infers execution order and parallelism from data dependencies — there is no explicit scheduler to write.
18
- - **Write once, run anywhere**: the same pipeline runs locally, on HPC (SLURM, SGE, LSF, PBS), and on cloud (AWS Batch, Google Batch, Azure Batch, Kubernetes) by changing config/profiles, not code.
19
- - **Reproducibility**: per-task containers (Docker/Singularity/Apptainer/Conda/Wave) + `-resume` caching + pinned pipeline revisions.
20
- - **DSL2** is the modern, required syntax: modular `process`/`workflow`/`include` definitions.
21
-
22
- This skill covers both **running** existing pipelines and **developing** your own (Nextflow language + nf-core conventions, testing with nf-test, configuration, and deployment).
23
-
24
- ## When to Use This Skill
25
-
26
- Use this skill when the user wants to:
27
- - Run an nf-core or custom Nextflow pipeline, or debug a failing/resuming run.
28
- - Write or modify `.nf` scripts, `nextflow.config`, profiles, or `nextflow_schema.json`.
29
- - Author or test nf-core-style modules/subworkflows (`main.nf`, `meta.yml`, `tests/`, nf-test).
30
- - Configure executors, containers, or resources; scale to HPC or cloud.
31
- - Build a reproducible scientific/bioinformatics workflow (even if "Nextflow" is not named).
32
- - Understand processes, channels, operators, `take`/`emit`, `publishDir`, `ext.args`, meta maps.
33
-
34
- ## Setup
35
-
36
- Nextflow needs **Bash** and **Java 17 or newer** (17–25 supported). Verify with `java -version`.
37
-
38
- ```bash
39
- # Install Nextflow (self-contained launcher)
40
- curl -s https://get.nextflow.io | bash # creates ./nextflow
41
- sudo mv nextflow /usr/local/bin/ # put on PATH
42
- nextflow info # verify
43
-
44
- # Or via conda/bioconda (also gets a managed Java)
45
- conda create -n nf -c bioconda -c conda-forge nextflow nf-core
46
- ```
47
-
48
- ```bash
49
- # nf-core tools (Python) for creating/linting/running nf-core assets
50
- uv pip install nf-core # or: conda install -c bioconda nf-core
51
- nf-core --version
52
- ```
53
-
54
- Pin the engine for reproducibility: `export NXF_VER=24.10.0` (use an [edge] release only if needed). For air-gapped/HPC, see `references/running-pipelines.md` (offline mode) and `references/configuration.md`.
55
-
56
- ## Two Modes of Work
57
-
58
- Decide which path the user is on — it changes everything:
59
-
60
- | Goal | Start here |
61
- |------|-----------|
62
- | **Run** an existing pipeline (nf-core or a `.nf` you were given) | `references/running-pipelines.md` |
63
- | **Develop** a new pipeline / module / subworkflow | `references/language.md` + `references/developing.md` |
64
- | **Configure / scale** (HPC, cloud, containers, resources) | `references/configuration.md` + `references/containers.md` |
65
- | **Test** modules/pipelines | `references/testing.md` |
66
-
67
- ## Quick Start
68
-
69
- ### Run an nf-core pipeline
70
-
71
- Always smoke-test with the bundled `test` profile first; it uses tiny data and proves your environment works.
72
-
73
- ```bash
74
- # 1. Confirm setup works (downloads pipeline + tiny test data)
75
- nextflow run nf-core/rnaseq -profile test,docker --outdir results
76
-
77
- # 2. Real run: pin a revision (-r), pick a container engine, pass inputs
78
- nextflow run nf-core/rnaseq -r 3.14.0 \
79
- -profile docker \
80
- --input samplesheet.csv \
81
- --genome GRCh38 \
82
- --outdir results \
83
- -resume
84
- ```
85
-
86
- - `-profile` (single dash) selects bundled config profiles; **combine** them comma-separated, e.g. `test,docker`. Container/infra profiles (`docker`, `singularity`, `conda`) are mutually exclusive — pick one.
87
- - `--input`, `--genome`, `--outdir` (double dash) are **pipeline** parameters. nf-core pipelines take a **samplesheet CSV**, not loose files.
88
- - `-resume` reuses cached results from the last run. `-r <version>` pins a release for reproducibility.
89
-
90
- Use `nf-core pipelines launch <name>` for an interactive, schema-validated way to build the command and a `-params-file`. See `references/running-pipelines.md`.
91
-
92
- ### Write a minimal pipeline
93
-
94
- ```nextflow
95
- #!/usr/bin/env nextflow
96
-
97
- process SAYHELLO {
98
- tag "$greeting"
99
- publishDir "results", mode: 'copy'
100
-
101
- input:
102
- val greeting
103
-
104
- output:
105
- path "${greeting}.txt"
106
-
107
- script:
108
- """
109
- echo '$greeting world' > ${greeting}.txt
110
- """
111
- }
112
-
113
- workflow {
114
- channel.of('hello', 'bonjour', 'hola') | SAYHELLO
115
- }
116
- ```
117
-
118
- ```bash
119
- nextflow run main.nf # add -resume on reruns
120
- ```
121
-
122
- The full language (processes, channels, operators, DSL2 workflows with `take`/`main`/`emit`, modules) is in `references/language.md`.
123
-
124
- ## Core Concepts at a Glance
125
-
126
- - **Process**: a unit of work that runs a script (Bash by default). Declares `input:`, `output:`, optional `directives` (resources, container, `publishDir`, `tag`, `errorStrategy`), and a `script:`/`shell:`/`exec:` block. Each task runs in its own isolated work directory (`work/xx/yy…`).
127
- - **Channel**: the async queues that connect processes. **Queue channels** are consumable streams; **value channels** hold a single reusable value. Created with factories like `channel.of`, `channel.fromPath`, `channel.fromFilePairs`, `channel.value`.
128
- - **Operator**: transforms/combines channels — `map`, `filter`, `collect`, `groupTuple`, `join`, `combine`, `mix`, `flatten`, `branch`, `multiMap`, `splitCsv`, `view`, `set`.
129
- - **Workflow**: composes processes. DSL2 workflows can declare `take:` (inputs), `main:` (logic), `emit:` (named outputs) and be `include`d as subworkflows. The unnamed `workflow {}` is the entry point.
130
- - **Module**: a `.nf` file exposing processes/workflows via `include { NAME } from './path'` (supports `as` aliasing).
131
- - **Configuration**: `nextflow.config` sets `params`, `process` directives, `executor`, container engines, and named `profiles`. Selectors `withName:`/`withLabel:` target specific processes. See `references/configuration.md`.
132
- - **meta map** (nf-core): the convention of carrying a metadata map (`[ id:'sample1', single_end:false ]`) alongside files in input/output tuples so samples stay labeled through the pipeline. See `references/developing.md`.
133
-
134
- ## nf-core tools CLI
135
-
136
- nf-core tools (v3+) group subcommands under `pipelines`, `modules`, and `subworkflows`. (Bare forms like `nf-core lint` still work but warn — prefer the grouped form.)
137
-
138
- | Command | Purpose |
139
- |---------|---------|
140
- | `nf-core pipelines list` | List/search nf-core pipelines (`--json`, keywords) |
141
- | `nf-core pipelines create` | Scaffold a new pipeline from the nf-core template |
142
- | `nf-core pipelines launch <name>` | Interactive, schema-driven run command + params file |
143
- | `nf-core pipelines download <name>` | Download pipeline + containers for offline/HPC use |
144
- | `nf-core pipelines lint` | Lint a pipeline against nf-core standards (run in repo root) |
145
- | `nf-core pipelines schema build` | Build/edit `nextflow_schema.json` via web GUI |
146
- | `nf-core pipelines create-params-file <name>` | Generate a documented YAML params file |
147
- | `nf-core pipelines bump-version` / `sync` | Bump version / sync with template updates |
148
- | `nf-core modules list/info/install/update/remove` | Manage modules from nf-core/modules |
149
- | `nf-core modules create` / `lint` / `test` | Author, lint, and nf-test a module |
150
- | `nf-core modules patch` / `bump-versions` | Patch an installed module / bump tool versions |
151
- | `nf-core subworkflows install/create/lint/test` | Same lifecycle for subworkflows |
152
-
153
- Full command reference, flags, and examples: `references/nf-core-tools.md`.
154
-
155
- ## Essential `nextflow` CLI
156
-
157
- | Command | Purpose |
158
- |---------|---------|
159
- | `nextflow run <pipeline> -profile <p> --outdir <dir>` | Run a pipeline (path, `.nf`, or `user/repo`) |
160
- | `-resume` | Reuse cached results from prior run |
161
- | `-r <rev>` | Run a specific git revision/tag/branch |
162
- | `-params-file params.yml` | Supply parameters from YAML/JSON |
163
- | `-c custom.config` | Layer in an extra config file |
164
- | `-with-report -with-trace -with-timeline -with-dag flow.html` | Execution report, trace, timeline, DAG |
165
- | `-stub-run` | Run `stub:` blocks only (dry-run plumbing) |
166
- | `nextflow log` | Inspect past runs |
167
- | `nextflow clean -f -before <run>` | Delete old `work/` data |
168
- | `nextflow pull / drop / list / info <repo>` | Manage cached remote pipelines |
169
-
170
- Config, executors, caching internals, and tracing details: `references/configuration.md`.
171
-
172
- ## Best Practices (high-value habits)
173
-
174
- - **Always `test` first**: `-profile test,docker` (or `singularity`/`conda`) before real data — fast and catches environment problems.
175
- - **Pin everything**: pipeline revision (`-r`), `NXF_VER`, and tool versions (containers). Don't run `latest` for science you'll publish.
176
- - **Use `-resume`** and understand caching: a task re-runs if its inputs, script, or container change. See cache-debugging in `references/configuration.md`.
177
- - **Parameterize via config/params-file**, not hardcoded paths. Keep `params` and profiles in `nextflow.config`.
178
- - **One container/conda env per process**; never rely on tools installed on the host.
179
- - **For nf-core dev**: reuse existing modules (`nf-core modules install`) before writing new ones; pass tool flags through `ext.args` (not hardcoded in the script); always include a `stub:` block and nf-test tests; run `nf-core pipelines lint` and `prettier` before committing.
180
- - **Right-size resources** with `process_low/medium/high` labels and `errorStrategy 'retry'` with dynamic `task.attempt` scaling instead of one giant request.
181
- - **Write forward-compatible syntax**: the strict-syntax parser becomes the default in Nextflow 26.04. Prefer lowercase `channel.of(...)`, explicit closure params (`{ v -> ... }`), `def` for all variables, and `emit:`-named outputs. Check with `nextflow lint`.
182
-
183
- ## Reference Files
184
-
185
- Read the relevant file when you need depth — each is self-contained:
186
-
187
- - `references/language.md` — DSL2 language: processes, directives, channels, operators, workflows (`take`/`emit`), modules, dynamic resources, error handling.
188
- - `references/configuration.md` — `nextflow.config`, scopes, `profiles`, `withName`/`withLabel` selectors, executors (local/SLURM/cloud), caching/`-resume` internals, tracing/reports, the `nextflow` CLI.
189
- - `references/containers.md` — Docker, Singularity/Apptainer, Podman, Conda, Wave containers; choosing and enabling engines; common gotchas.
190
- - `references/running-pipelines.md` — finding/running nf-core pipelines, samplesheets, params files, reference genomes (iGenomes), offline runs, institutional configs, Seqera Platform.
191
- - `references/nf-core-tools.md` — complete `nf-core` CLI reference (pipelines/modules/subworkflows), flags, and workflows.
192
- - `references/developing.md` — authoring nf-core pipelines & modules: template layout, module `main.nf`/`meta.yml`, meta maps, `ext.args`/`modules.config`, subworkflows, resource labels, linting & Harshil alignment style.
193
- - `references/testing.md` — nf-test for modules/subworkflows/pipelines: test structure, assertions, snapshots, tags, running tests, CI.
194
-
195
- Official docs: Nextflow https://www.nextflow.io/docs/latest/ · nf-core https://nf-co.re/docs/ · Training https://training.nextflow.io/