How VoiceStudio Uses Pyannote and WhisperX for Speaker Diarization: Complete Pipeline Guide

VoiceStudio combines a Pyannote audio diarization model with the WhisperX ASR engine in a two-stage pipeline, then merges their outputs through time-weighted overlap matching and segment re-splitting to produce speaker-attributed transcripts. The integrated results power downstream speaker cloning, subtitle exports, and API delivery.

Speaker diarization—the task of answering "who spoke when?"—is a core capability of the open-source VoiceStudio project. This article breaks down exactly how the repository leverages Pyannote for speaker boundary detection and WhisperX for transcription, then explains the fusion logic that unifies both streams. All technical details are derived directly from the source code implementation.


Loading the Pyannote Diarization Pipeline

The diarization model is managed through lazy initialization in backend/services/model_manager.py. The function get_diarization_pipeline() handles the complete setup lifecycle.


# backend/services/model_manager.py (lines 3465-3510)

def get_diarization_pipeline():
    """
    Lazily loads and caches the pretrained Pyannote speaker-diarization-3.1 model.
    Handles Hugging Face Hub token compatibility and safe-globals registration.
    """

Key responsibilities of this loader:

  • Model specification: Downloads pyannote/speaker-diarization-3.1 from Hugging Face Hub
  • Token authentication: Manages HF Hub access tokens for gated model access
  • Safe globals registration: Configures torch.serialization for secure checkpoint loading
  • Pipeline caching: Returns a cached instance after first call to avoid repeated model loading

The pipeline object returned is a ready-to-use Pyannote Pipeline that accepts audio tensors and returns speaker turn annotations with start/end timestamps and speaker labels.


Running WhisperX for Time-Aligned Transcription

WhisperX serves as the default ASR backend in backend/services/asr_backend.py. Unlike standard Whisper, WhisperX adds forced alignment for precise word-level timestamps—critical for accurate speaker assignment.


# backend/services/asr_backend.py (lines 500-560)

import whisperx

# Load the ASR model

model, metadata = whisperx.load_model("large-v2", device="cuda")

# Obtain aligned transcription with word-level timestamps

result = whisperx.align(audio, model, metadata)

WhisperX output format:

Field Description
segments List of transcript segments with start, end, text
word_segments Individual words with sub-second timing precision
language Detected or specified language code

The alignment step is what makes downstream fusion possible—without accurate timestamps, speaker assignment would fail.


Merging Diarization and Transcription Results

The core integration logic lives in backend/services/segmentation.py. Two complementary functions handle the fusion: assign_speakers_from_diarization() and resplit_segments_by_diarization().

Speaker Assignment via Time-Overlap Weighting

assign_speakers_from_diarization() (lines 558-595) replaces WhisperX's placeholder speaker IDs with actual Pyannote speaker labels:


# backend/services/segmentation.py

def assign_speakers_from_diarization(asr_segments, diar_pipeline):
    """
    Assigns each transcript segment to the Pyannote speaker 
    with maximum temporal overlap.
    """
    for segment in asr_segments:
        # Calculate weighted overlap with each Pyannote speaker turn

        overlaps = compute_overlap_weights(
            segment["start"], 
            segment["end"],
            diar_pipeline(audio)  # speaker turns with (start, end, speaker)

        )
        segment["speaker"] = argmax(overlaps)  # most overlapping speaker

    return asr_segments

Overlap-weighted strategy explained:

  • For each WhisperX transcript segment, compute temporal intersection with every Pyannote speaker turn
  • Weight by duration of overlap, not just binary intersection
  • Assign to the speaker with highest cumulative overlap score

This handles edge cases where segment boundaries don't perfectly align with speaker changes.

Segment Re-Splitting at Speaker Boundaries

resplit_segments_by_diarization() (lines 785-820) performs finer-grained segmentation:


# backend/services/segmentation.py

def resplit_segments_by_diarization(segments, diar_pipeline):
    """
    Re-splits transcript segments when Pyannote detects speaker changes
    mid-segment, preserving WhisperX word-level timing.
    """
    for segment in segments:
        speaker_turns = diar_pipeline.get_turns(
            segment["start"], 
            segment["end"]
        )
        if len(speaker_turns) > 1:
            # Split this segment at speaker change points

            sub_segments = split_at_boundaries(
                segment, 
                [t.end for t in speaker_turns[:-1]]
            )
            yield from sub_segments
        else:
            yield segment

Why re-splitting matters:

  • WhisperX segments may span multiple speaker turns
  • Clean subtitle files require speaker changes at precise timestamps
  • Word-level timing from WhisperX is preserved through the split operation

Downstream Integration: Where Diarized Output Flows

The unified speaker-attributed transcript ({speaker, start, end, text}) feeds multiple downstream systems:

Speaker Cloning Service

backend/services/speaker_clone.py (lines 80-110) maps diarized speaker IDs to voice models:


# Map detected speaker labels to clone models

speaker_id = segment["speaker"]  # e.g., "SPEAKER_01"

clone_model = speaker_registry.get_or_create(speaker_id)

Each unique speaker detected by Pyannote becomes a target for voice cloning operations.

Export Formats and API Delivery

The final segments render to multiple output formats:

  • SRT/VTT: Speaker labels prepended to subtitle text ([SPEAKER_01] Hello world)
  • JSON: Structured array with full metadata for UI consumption
  • API response: Direct passthrough to frontend for real-time display

Complete Pipeline Example

from services.model_manager import get_diarization_pipeline
from services.segmentation import (
    assign_speakers_from_diarization,
    resplit_segments_by_diarization
)
import whisperx

# 1. Load models (cached after first call)

diar_pipeline = get_diarization_pipeline()
asr_model, metadata = whisperx.load_model("large-v2")

# 2. Process audio

audio = load_audio("meeting.wav")
diar_result = diar_pipeline(audio)  # Pyannote speaker turns

asr_result = whisperx.align(audio, asr_model, metadata)

# 3. Fuse results

segments = assign_speakers_from_diarization(
    asr_result["segments"], 
    diar_result
)
final_segments = list(resplit_segments_by_diarization(
    segments, 
    diar_result
))

# 4. Consume downstream

for seg in final_segments:
    print(f"[{seg['speaker']}] {seg['start']:.2f}-{seg['end']:.2f}: {seg['text']}")

Summary

  • Pyannote (pyannote/speaker-diarization-3.1) provides speaker boundary detection via lazily-loaded pipeline in model_manager.py
  • WhisperX delivers high-accuracy transcription with word-level timestamps through align() in the ASR backend
  • Fusion logic in segmentation.py uses time-overlap weighting for speaker assignment and re-splitting for clean segment boundaries
  • Output integration feeds speaker cloning, multi-format exports, and API delivery with unified {speaker, start, end, text} structures
  • Test coverage in tests/test_assign_speakers_from_diarization.py and tests/test_resplit_speaker.py validates edge cases

Frequently Asked Questions

What version of Pyannote does VoiceStudio use?

VoiceStudio specifically uses pyannote/speaker-diarization-3.1, the current state-of-the-art release from Hugging Face Hub. The get_diarization_pipeline() function in backend/services/model_manager.py pins this model identifier and handles authentication for the gated repository.

Why use WhisperX instead of standard OpenAI Whisper?

WhisperX adds forced phonetic alignment that produces sub-second word-level timestamps. Standard Whisper only provides segment-level timing (~10-30 second chunks), which is too coarse for accurate speaker assignment when diarization boundaries fall mid-segment. The whisperx.align() call in backend/services/asr_backend.py delivers the precision needed for overlap-based fusion.

How does VoiceStudio handle overlapping speech?

The assign_speakers_from_diarization() function uses weighted temporal overlap rather than simple binary intersection. When multiple speakers overlap a transcript segment, the speaker with the greatest cumulative overlap time wins assignment. For complete overlap scenarios, the dominant speaker is annotated; advanced multi-label assignment would require extending the segmentation logic.

Can the diarization pipeline run on CPU only?

Yes, though with significant latency tradeoffs. The get_diarization_pipeline() function respects PyTorch device configuration—CUDA is preferred when available, but CPU fallback is supported. For production deployments, GPU acceleration is strongly recommended; the Pyannote diarization-3.1 model contains speaker embedding networks that benefit substantially from parallel computation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →