How to Use the Neuropixels-Analysis Toolkit for Spike Sorting Quality Control

The Neuropixels-Analysis toolkit provides a modular Python pipeline that automates spike sorting quality control through preprocessing, motion correction, metric computation, and AI-assisted curation.

The Neuropixels-Analysis skill in the K-Dense-AI/scientific-agent-skills repository wraps the full SpikeInterface ecosystem to deliver rigorous quality control for high-density electrophysiology recordings. Its architecture decomposes the workflow into discrete, composable stages—from raw data ingestion to publish-ready reporting—that can run independently or chain together via the npa.run_pipeline() convenience function.

Core Pipeline Architecture

The toolkit implements a seven-layer architecture documented in SKILL.md and references/SPIKE_SORTING.md:

  • Data Ingestion: Reads SpikeGLX, Open-Ephys, or NWB files via si.read_spikeglx(), si.read_openephys(), or si.read_nwb() to create a RecordingExtractor object.

  • Preprocessing: Applies high-pass filtering, phase-shift correction (for Neuropixels 1.0 probes), bad-channel detection, and common-average referencing using si.highpass_filter(), si.phase_shift(), and si.common_reference().

  • Motion Correction: Estimates longitudinal drift via npa.estimate_motion() with two presets—a fast kilosort-like mode and a non-rigid accurate mode—and applies correction through npa.correct_motion().

  • Spike Sorting: Executes GPU-accelerated Kilosort4 by default via si.run_sorter(), with fallback support for Spyking-Circus2, Mountainsort5, and other SpikeInterface-compatible algorithms.

  • Post-processing: Builds a SortingAnalyzer via si.create_sorting_analyzer() and computes waveforms, templates, amplitudes, correlograms, and unit locations.

  • Metric Computation: Calculates the full suite of quality metrics including SNR, ISI violations, presence ratio, amplitude cutoff, and drift metrics via analyzer.compute('quality_metrics').

  • Curation: Applies Allen Institute or IBL-derived thresholds through npa.curate(), with optional AI-assisted review via npa.analyze_unit_visually() for borderline units.

Loading Raw Neuropixels Recordings

Begin by importing the toolkit and configuring parallel processing. The job_kwargs dictionary controls CPU utilization and chunking strategy:

import spikeinterface.full as si
import neuropixels_analysis as npa

# Configure parallel processing

job_kwargs = dict(n_jobs=-1, chunk_duration="1s", progress_bar=True)

# Load SpikeGLX data (adjust stream_id for your probe)

recording = si.read_spikeglx("/data/neuropixels/recording_g0", stream_id="imec0.ap")

# Optional: Slice first 60 seconds for rapid testing

recording = recording.frame_slice(0, int(60 * recording.get_sampling_frequency()))

Preprocessing and Drift Correction

Execute the full preprocessing chain through npa.preprocess(), which implements the standard Neuropixels pipeline including high-pass filtering and phase-shift correction:


# Apply high-pass filter, phase shift, and common average reference

rec = npa.preprocess(recording)

# Estimate motion using the fast kilosort-like preset

motion_info = npa.estimate_motion(rec, preset="kilosort_like")
print(f"Maximum drift detected: {motion_info['motion'].max():.1f} µm")

# Visualize drift map

npa.plot_drift(rec, motion_info, output="drift_map.png")

If drift exceeds approximately 10 µm, apply the more accurate non-rigid correction:

if motion_info["motion"].max() > 10:
    rec = npa.correct_motion(rec, preset="nonrigid_accurate")

Running Spike Sorting and Computing Quality Metrics

Execute spike sorting and build the SortingAnalyzer to extract quality metrics. According to scripts/run_sorting.py, Kilosort4 serves as the default sorter for GPU-accelerated processing:


# Run Kilosort4 (requires GPU; substitute 'tridesclous2' or 'spykingcircus2' for CPU-only)

sorting = si.run_sorter("kilosort4", rec, folder="ks4_output", **job_kwargs)

# Create analyzer and compute quality metrics

analyzer = si.create_sorting_analyzer(sorting, rec, sparse=True)
analyzer.compute("quality_metrics")

# Extract metrics DataFrame

metrics = analyzer.get_extension("quality_metrics").get_data()
print(metrics.head())

The metrics DataFrame contains standard quality indicators including firing rate, presence ratio, ISI violations ratio, amplitude cutoff, and signal-to-noise ratio.

Automated and AI-Assisted Curation

Apply rule-based curation using Allen Institute thresholds via npa.curate():


# Automated labeling based on Allen criteria

labels = npa.curate(metrics, method="allen")
good_units = labels["good"]

For borderline units with ambiguous metrics, invoke AI-assisted visual analysis through npa.analyze_unit_visually():

from anthropic import Anthropic

client = Anthropic()  # API key via environment variable

uncertain_units = metrics.query("snr > 3 and snr < 8").index.tolist()

for unit_id in uncertain_units:
    result = npa.analyze_unit_visually(analyzer, unit_id, api_client=client)
    print(f"Unit {unit_id}: {result['classification']}")
    print(f"Reasoning: {result['reasoning'][:120]}...")

When executing within Claude Code, the API client injects automatically and explicit instantiation can be omitted.

Generating Reports and Exporting Results

Create publication-ready deliverables using npa.generate_analysis_report() and standard SpikeInterface export functions:


# Generate HTML report with plots and metrics

report_dir = npa.generate_analysis_report(
    {"sorting": sorting, "metrics": metrics, "analyzer": analyzer},
    output_dir="report/",
)

# Export to Phy for manual review

si.export_to_phy(analyzer, output_folder="phy_export/", compute_pc_features=True)

# Save metrics table for downstream analysis

metrics.to_csv("quality_metrics.csv")

Reference implementations for these workflows reside in scripts/compute_metrics.py and assets/analysis_template.py.

Summary

  • The Neuropixels-Analysis toolkit wraps SpikeInterface into a modular pipeline specifically designed for spike sorting quality control on Neuropixels recordings.
  • Key functions include npa.preprocess() for signal conditioning, npa.estimate_motion() for drift quantification, and analyzer.compute('quality_metrics') for validation metrics.
  • Automated curation applies Allen Institute criteria via npa.curate(), while npa.analyze_unit_visually() provides AI-assisted classification for ambiguous units.
  • Complete workflows are documented in SKILL.md and implemented in scripts/run_sorting.py and assets/analysis_template.py.

Frequently Asked Questions

What file formats does the Neuropixels-Analysis toolkit support?

The toolkit supports SpikeGLX (via si.read_spikeglx()), Open-Ephys (via si.read_openephys()), and NWB (via si.read_nwb()), creating standardized RecordingExtractor objects regardless of source format.

How does the toolkit handle motion correction for drifting units?

The npa.estimate_motion() function quantifies longitudinal drift with two presets: a fast kilosort-like mode for quick estimates and a nonrigid_accurate preset for precise correction when drift exceeds 10 µm, implemented through npa.correct_motion().

Can I use the toolkit without a GPU for spike sorting?

Yes. While Kilosort4 serves as the default GPU-accelerated sorter, the si.run_sorter() function accepts CPU-only alternatives including "tridesclous2" and "spykingcircus2" specified via the first argument.

What quality metrics are computed for spike sorting validation?

The analyzer.compute('quality_metrics') extension calculates SNR, ISI violations ratio, presence ratio, amplitude cutoff, firing rate, and drift metrics, returning a pandas DataFrame suitable for automated thresholding or manual review.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →