How to Use the Neuropixels-Analysis Toolkit for Spike Sorting Quality Control
The Neuropixels-Analysis toolkit provides a modular Python pipeline that automates spike sorting quality control through preprocessing, motion correction, metric computation, and AI-assisted curation.
The Neuropixels-Analysis skill in the K-Dense-AI/scientific-agent-skills repository wraps the full SpikeInterface ecosystem to deliver rigorous quality control for high-density electrophysiology recordings. Its architecture decomposes the workflow into discrete, composable stages—from raw data ingestion to publish-ready reporting—that can run independently or chain together via the npa.run_pipeline() convenience function.
Core Pipeline Architecture
The toolkit implements a seven-layer architecture documented in SKILL.md and references/SPIKE_SORTING.md:
-
Data Ingestion: Reads SpikeGLX, Open-Ephys, or NWB files via
si.read_spikeglx(),si.read_openephys(), orsi.read_nwb()to create aRecordingExtractorobject. -
Preprocessing: Applies high-pass filtering, phase-shift correction (for Neuropixels 1.0 probes), bad-channel detection, and common-average referencing using
si.highpass_filter(),si.phase_shift(), andsi.common_reference(). -
Motion Correction: Estimates longitudinal drift via
npa.estimate_motion()with two presets—a fast kilosort-like mode and a non-rigid accurate mode—and applies correction throughnpa.correct_motion(). -
Spike Sorting: Executes GPU-accelerated Kilosort4 by default via
si.run_sorter(), with fallback support for Spyking-Circus2, Mountainsort5, and other SpikeInterface-compatible algorithms. -
Post-processing: Builds a
SortingAnalyzerviasi.create_sorting_analyzer()and computes waveforms, templates, amplitudes, correlograms, and unit locations. -
Metric Computation: Calculates the full suite of quality metrics including SNR, ISI violations, presence ratio, amplitude cutoff, and drift metrics via
analyzer.compute('quality_metrics'). -
Curation: Applies Allen Institute or IBL-derived thresholds through
npa.curate(), with optional AI-assisted review vianpa.analyze_unit_visually()for borderline units.
Loading Raw Neuropixels Recordings
Begin by importing the toolkit and configuring parallel processing. The job_kwargs dictionary controls CPU utilization and chunking strategy:
import spikeinterface.full as si
import neuropixels_analysis as npa
# Configure parallel processing
job_kwargs = dict(n_jobs=-1, chunk_duration="1s", progress_bar=True)
# Load SpikeGLX data (adjust stream_id for your probe)
recording = si.read_spikeglx("/data/neuropixels/recording_g0", stream_id="imec0.ap")
# Optional: Slice first 60 seconds for rapid testing
recording = recording.frame_slice(0, int(60 * recording.get_sampling_frequency()))
Preprocessing and Drift Correction
Execute the full preprocessing chain through npa.preprocess(), which implements the standard Neuropixels pipeline including high-pass filtering and phase-shift correction:
# Apply high-pass filter, phase shift, and common average reference
rec = npa.preprocess(recording)
# Estimate motion using the fast kilosort-like preset
motion_info = npa.estimate_motion(rec, preset="kilosort_like")
print(f"Maximum drift detected: {motion_info['motion'].max():.1f} µm")
# Visualize drift map
npa.plot_drift(rec, motion_info, output="drift_map.png")
If drift exceeds approximately 10 µm, apply the more accurate non-rigid correction:
if motion_info["motion"].max() > 10:
rec = npa.correct_motion(rec, preset="nonrigid_accurate")
Running Spike Sorting and Computing Quality Metrics
Execute spike sorting and build the SortingAnalyzer to extract quality metrics. According to scripts/run_sorting.py, Kilosort4 serves as the default sorter for GPU-accelerated processing:
# Run Kilosort4 (requires GPU; substitute 'tridesclous2' or 'spykingcircus2' for CPU-only)
sorting = si.run_sorter("kilosort4", rec, folder="ks4_output", **job_kwargs)
# Create analyzer and compute quality metrics
analyzer = si.create_sorting_analyzer(sorting, rec, sparse=True)
analyzer.compute("quality_metrics")
# Extract metrics DataFrame
metrics = analyzer.get_extension("quality_metrics").get_data()
print(metrics.head())
The metrics DataFrame contains standard quality indicators including firing rate, presence ratio, ISI violations ratio, amplitude cutoff, and signal-to-noise ratio.
Automated and AI-Assisted Curation
Apply rule-based curation using Allen Institute thresholds via npa.curate():
# Automated labeling based on Allen criteria
labels = npa.curate(metrics, method="allen")
good_units = labels["good"]
For borderline units with ambiguous metrics, invoke AI-assisted visual analysis through npa.analyze_unit_visually():
from anthropic import Anthropic
client = Anthropic() # API key via environment variable
uncertain_units = metrics.query("snr > 3 and snr < 8").index.tolist()
for unit_id in uncertain_units:
result = npa.analyze_unit_visually(analyzer, unit_id, api_client=client)
print(f"Unit {unit_id}: {result['classification']}")
print(f"Reasoning: {result['reasoning'][:120]}...")
When executing within Claude Code, the API client injects automatically and explicit instantiation can be omitted.
Generating Reports and Exporting Results
Create publication-ready deliverables using npa.generate_analysis_report() and standard SpikeInterface export functions:
# Generate HTML report with plots and metrics
report_dir = npa.generate_analysis_report(
{"sorting": sorting, "metrics": metrics, "analyzer": analyzer},
output_dir="report/",
)
# Export to Phy for manual review
si.export_to_phy(analyzer, output_folder="phy_export/", compute_pc_features=True)
# Save metrics table for downstream analysis
metrics.to_csv("quality_metrics.csv")
Reference implementations for these workflows reside in scripts/compute_metrics.py and assets/analysis_template.py.
Summary
- The Neuropixels-Analysis toolkit wraps SpikeInterface into a modular pipeline specifically designed for spike sorting quality control on Neuropixels recordings.
- Key functions include
npa.preprocess()for signal conditioning,npa.estimate_motion()for drift quantification, andanalyzer.compute('quality_metrics')for validation metrics. - Automated curation applies Allen Institute criteria via
npa.curate(), whilenpa.analyze_unit_visually()provides AI-assisted classification for ambiguous units. - Complete workflows are documented in
SKILL.mdand implemented inscripts/run_sorting.pyandassets/analysis_template.py.
Frequently Asked Questions
What file formats does the Neuropixels-Analysis toolkit support?
The toolkit supports SpikeGLX (via si.read_spikeglx()), Open-Ephys (via si.read_openephys()), and NWB (via si.read_nwb()), creating standardized RecordingExtractor objects regardless of source format.
How does the toolkit handle motion correction for drifting units?
The npa.estimate_motion() function quantifies longitudinal drift with two presets: a fast kilosort-like mode for quick estimates and a nonrigid_accurate preset for precise correction when drift exceeds 10 µm, implemented through npa.correct_motion().
Can I use the toolkit without a GPU for spike sorting?
Yes. While Kilosort4 serves as the default GPU-accelerated sorter, the si.run_sorter() function accepts CPU-only alternatives including "tridesclous2" and "spykingcircus2" specified via the first argument.
What quality metrics are computed for spike sorting validation?
The analyzer.compute('quality_metrics') extension calculates SNR, ISI violations ratio, presence ratio, amplitude cutoff, firing rate, and drift metrics, returning a pandas DataFrame suitable for automated thresholding or manual review.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →