# How to Use the Neuropixels-Analysis Toolkit for Spike Sorting Quality Control

> Master spike sorting quality control with Neuropixels-Analysis. Automate preprocessing, motion correction, and AI curation with this powerful Python toolkit.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**The Neuropixels-Analysis toolkit provides a modular Python pipeline that automates spike sorting quality control through preprocessing, motion correction, metric computation, and AI-assisted curation.**

The **Neuropixels-Analysis** skill in the `K-Dense-AI/scientific-agent-skills` repository wraps the full SpikeInterface ecosystem to deliver rigorous quality control for high-density electrophysiology recordings. Its architecture decomposes the workflow into discrete, composable stages—from raw data ingestion to publish-ready reporting—that can run independently or chain together via the `npa.run_pipeline()` convenience function.

## Core Pipeline Architecture

The toolkit implements a seven-layer architecture documented in [`SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/SKILL.md) and [`references/SPIKE_SORTING.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/references/SPIKE_SORTING.md):

- **Data Ingestion**: Reads **SpikeGLX**, **Open-Ephys**, or **NWB** files via `si.read_spikeglx()`, `si.read_openephys()`, or `si.read_nwb()` to create a `RecordingExtractor` object.

- **Preprocessing**: Applies high-pass filtering, phase-shift correction (for Neuropixels 1.0 probes), bad-channel detection, and common-average referencing using `si.highpass_filter()`, `si.phase_shift()`, and `si.common_reference()`.

- **Motion Correction**: Estimates longitudinal drift via `npa.estimate_motion()` with two presets—a fast *kilosort-like* mode and a non-rigid accurate mode—and applies correction through `npa.correct_motion()`.

- **Spike Sorting**: Executes GPU-accelerated **Kilosort4** by default via `si.run_sorter()`, with fallback support for **Spyking-Circus2**, **Mountainsort5**, and other SpikeInterface-compatible algorithms.

- **Post-processing**: Builds a `SortingAnalyzer` via `si.create_sorting_analyzer()` and computes waveforms, templates, amplitudes, correlograms, and unit locations.

- **Metric Computation**: Calculates the full suite of quality metrics including **SNR**, **ISI violations**, **presence ratio**, **amplitude cutoff**, and drift metrics via `analyzer.compute('quality_metrics')`.

- **Curation**: Applies Allen Institute or IBL-derived thresholds through `npa.curate()`, with optional AI-assisted review via `npa.analyze_unit_visually()` for borderline units.

## Loading Raw Neuropixels Recordings

Begin by importing the toolkit and configuring parallel processing. The `job_kwargs` dictionary controls CPU utilization and chunking strategy:

```python
import spikeinterface.full as si
import neuropixels_analysis as npa

# Configure parallel processing

job_kwargs = dict(n_jobs=-1, chunk_duration="1s", progress_bar=True)

# Load SpikeGLX data (adjust stream_id for your probe)

recording = si.read_spikeglx("/data/neuropixels/recording_g0", stream_id="imec0.ap")

# Optional: Slice first 60 seconds for rapid testing

recording = recording.frame_slice(0, int(60 * recording.get_sampling_frequency()))

```

## Preprocessing and Drift Correction

Execute the full preprocessing chain through `npa.preprocess()`, which implements the standard Neuropixels pipeline including high-pass filtering and phase-shift correction:

```python

# Apply high-pass filter, phase shift, and common average reference

rec = npa.preprocess(recording)

# Estimate motion using the fast kilosort-like preset

motion_info = npa.estimate_motion(rec, preset="kilosort_like")
print(f"Maximum drift detected: {motion_info['motion'].max():.1f} µm")

# Visualize drift map

npa.plot_drift(rec, motion_info, output="drift_map.png")

```

If drift exceeds approximately 10 µm, apply the more accurate non-rigid correction:

```python
if motion_info["motion"].max() > 10:
    rec = npa.correct_motion(rec, preset="nonrigid_accurate")

```

## Running Spike Sorting and Computing Quality Metrics

Execute spike sorting and build the `SortingAnalyzer` to extract quality metrics. According to [`scripts/run_sorting.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scripts/run_sorting.py), Kilosort4 serves as the default sorter for GPU-accelerated processing:

```python

# Run Kilosort4 (requires GPU; substitute 'tridesclous2' or 'spykingcircus2' for CPU-only)

sorting = si.run_sorter("kilosort4", rec, folder="ks4_output", **job_kwargs)

# Create analyzer and compute quality metrics

analyzer = si.create_sorting_analyzer(sorting, rec, sparse=True)
analyzer.compute("quality_metrics")

# Extract metrics DataFrame

metrics = analyzer.get_extension("quality_metrics").get_data()
print(metrics.head())

```

The metrics DataFrame contains standard quality indicators including firing rate, presence ratio, ISI violations ratio, amplitude cutoff, and signal-to-noise ratio.

## Automated and AI-Assisted Curation

Apply rule-based curation using Allen Institute thresholds via `npa.curate()`:

```python

# Automated labeling based on Allen criteria

labels = npa.curate(metrics, method="allen")
good_units = labels["good"]

```

For borderline units with ambiguous metrics, invoke AI-assisted visual analysis through `npa.analyze_unit_visually()`:

```python
from anthropic import Anthropic

client = Anthropic()  # API key via environment variable

uncertain_units = metrics.query("snr > 3 and snr < 8").index.tolist()

for unit_id in uncertain_units:
    result = npa.analyze_unit_visually(analyzer, unit_id, api_client=client)
    print(f"Unit {unit_id}: {result['classification']}")
    print(f"Reasoning: {result['reasoning'][:120]}...")

```

When executing within Claude Code, the API client injects automatically and explicit instantiation can be omitted.

## Generating Reports and Exporting Results

Create publication-ready deliverables using `npa.generate_analysis_report()` and standard SpikeInterface export functions:

```python

# Generate HTML report with plots and metrics

report_dir = npa.generate_analysis_report(
    {"sorting": sorting, "metrics": metrics, "analyzer": analyzer},
    output_dir="report/",
)

# Export to Phy for manual review

si.export_to_phy(analyzer, output_folder="phy_export/", compute_pc_features=True)

# Save metrics table for downstream analysis

metrics.to_csv("quality_metrics.csv")

```

Reference implementations for these workflows reside in [`scripts/compute_metrics.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scripts/compute_metrics.py) and [`assets/analysis_template.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/assets/analysis_template.py).

## Summary

- The **Neuropixels-Analysis toolkit** wraps SpikeInterface into a modular pipeline specifically designed for spike sorting quality control on Neuropixels recordings.
- **Key functions** include `npa.preprocess()` for signal conditioning, `npa.estimate_motion()` for drift quantification, and `analyzer.compute('quality_metrics')` for validation metrics.
- **Automated curation** applies Allen Institute criteria via `npa.curate()`, while `npa.analyze_unit_visually()` provides AI-assisted classification for ambiguous units.
- **Complete workflows** are documented in [`SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/SKILL.md) and implemented in [`scripts/run_sorting.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scripts/run_sorting.py) and [`assets/analysis_template.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/assets/analysis_template.py).

## Frequently Asked Questions

### What file formats does the Neuropixels-Analysis toolkit support?

The toolkit supports **SpikeGLX** (via `si.read_spikeglx()`), **Open-Ephys** (via `si.read_openephys()`), and **NWB** (via `si.read_nwb()`), creating standardized `RecordingExtractor` objects regardless of source format.

### How does the toolkit handle motion correction for drifting units?

The `npa.estimate_motion()` function quantifies longitudinal drift with two presets: a fast *kilosort-like* mode for quick estimates and a `nonrigid_accurate` preset for precise correction when drift exceeds 10 µm, implemented through `npa.correct_motion()`.

### Can I use the toolkit without a GPU for spike sorting?

Yes. While Kilosort4 serves as the default GPU-accelerated sorter, the `si.run_sorter()` function accepts CPU-only alternatives including **"tridesclous2"** and **"spykingcircus2"** specified via the first argument.

### What quality metrics are computed for spike sorting validation?

The `analyzer.compute('quality_metrics')` extension calculates **SNR**, **ISI violations ratio**, **presence ratio**, **amplitude cutoff**, **firing rate**, and drift metrics, returning a pandas DataFrame suitable for automated thresholding or manual review.