How VoiceStudio Uses Pyannote and WhisperX for Speaker Diarization: Complete Pipeline Guide
VoiceStudio combines a Pyannote audio diarization model with the WhisperX ASR engine in a two-stage pipeline, then merges their outputs through time-weighted overlap matching and segment re-splitting to produce speaker-attributed transcripts. The integrated results power downstream speaker cloning, subtitle exports, and API delivery.
Speaker diarization—the task of answering "who spoke when?"—is a core capability of the open-source VoiceStudio project. This article breaks down exactly how the repository leverages Pyannote for speaker boundary detection and WhisperX for transcription, then explains the fusion logic that unifies both streams. All technical details are derived directly from the source code implementation.
Loading the Pyannote Diarization Pipeline
The diarization model is managed through lazy initialization in backend/services/model_manager.py. The function get_diarization_pipeline() handles the complete setup lifecycle.
# backend/services/model_manager.py (lines 3465-3510)
def get_diarization_pipeline():
"""
Lazily loads and caches the pretrained Pyannote speaker-diarization-3.1 model.
Handles Hugging Face Hub token compatibility and safe-globals registration.
"""
Key responsibilities of this loader:
- Model specification: Downloads
pyannote/speaker-diarization-3.1from Hugging Face Hub - Token authentication: Manages HF Hub access tokens for gated model access
- Safe globals registration: Configures
torch.serializationfor secure checkpoint loading - Pipeline caching: Returns a cached instance after first call to avoid repeated model loading
The pipeline object returned is a ready-to-use Pyannote Pipeline that accepts audio tensors and returns speaker turn annotations with start/end timestamps and speaker labels.
Running WhisperX for Time-Aligned Transcription
WhisperX serves as the default ASR backend in backend/services/asr_backend.py. Unlike standard Whisper, WhisperX adds forced alignment for precise word-level timestamps—critical for accurate speaker assignment.
# backend/services/asr_backend.py (lines 500-560)
import whisperx
# Load the ASR model
model, metadata = whisperx.load_model("large-v2", device="cuda")
# Obtain aligned transcription with word-level timestamps
result = whisperx.align(audio, model, metadata)
WhisperX output format:
| Field | Description |
|---|---|
segments |
List of transcript segments with start, end, text |
word_segments |
Individual words with sub-second timing precision |
language |
Detected or specified language code |
The alignment step is what makes downstream fusion possible—without accurate timestamps, speaker assignment would fail.
Merging Diarization and Transcription Results
The core integration logic lives in backend/services/segmentation.py. Two complementary functions handle the fusion: assign_speakers_from_diarization() and resplit_segments_by_diarization().
Speaker Assignment via Time-Overlap Weighting
assign_speakers_from_diarization() (lines 558-595) replaces WhisperX's placeholder speaker IDs with actual Pyannote speaker labels:
# backend/services/segmentation.py
def assign_speakers_from_diarization(asr_segments, diar_pipeline):
"""
Assigns each transcript segment to the Pyannote speaker
with maximum temporal overlap.
"""
for segment in asr_segments:
# Calculate weighted overlap with each Pyannote speaker turn
overlaps = compute_overlap_weights(
segment["start"],
segment["end"],
diar_pipeline(audio) # speaker turns with (start, end, speaker)
)
segment["speaker"] = argmax(overlaps) # most overlapping speaker
return asr_segments
Overlap-weighted strategy explained:
- For each WhisperX transcript segment, compute temporal intersection with every Pyannote speaker turn
- Weight by duration of overlap, not just binary intersection
- Assign to the speaker with highest cumulative overlap score
This handles edge cases where segment boundaries don't perfectly align with speaker changes.
Segment Re-Splitting at Speaker Boundaries
resplit_segments_by_diarization() (lines 785-820) performs finer-grained segmentation:
# backend/services/segmentation.py
def resplit_segments_by_diarization(segments, diar_pipeline):
"""
Re-splits transcript segments when Pyannote detects speaker changes
mid-segment, preserving WhisperX word-level timing.
"""
for segment in segments:
speaker_turns = diar_pipeline.get_turns(
segment["start"],
segment["end"]
)
if len(speaker_turns) > 1:
# Split this segment at speaker change points
sub_segments = split_at_boundaries(
segment,
[t.end for t in speaker_turns[:-1]]
)
yield from sub_segments
else:
yield segment
Why re-splitting matters:
- WhisperX segments may span multiple speaker turns
- Clean subtitle files require speaker changes at precise timestamps
- Word-level timing from WhisperX is preserved through the split operation
Downstream Integration: Where Diarized Output Flows
The unified speaker-attributed transcript ({speaker, start, end, text}) feeds multiple downstream systems:
Speaker Cloning Service
backend/services/speaker_clone.py (lines 80-110) maps diarized speaker IDs to voice models:
# Map detected speaker labels to clone models
speaker_id = segment["speaker"] # e.g., "SPEAKER_01"
clone_model = speaker_registry.get_or_create(speaker_id)
Each unique speaker detected by Pyannote becomes a target for voice cloning operations.
Export Formats and API Delivery
The final segments render to multiple output formats:
- SRT/VTT: Speaker labels prepended to subtitle text (
[SPEAKER_01] Hello world) - JSON: Structured array with full metadata for UI consumption
- API response: Direct passthrough to frontend for real-time display
Complete Pipeline Example
from services.model_manager import get_diarization_pipeline
from services.segmentation import (
assign_speakers_from_diarization,
resplit_segments_by_diarization
)
import whisperx
# 1. Load models (cached after first call)
diar_pipeline = get_diarization_pipeline()
asr_model, metadata = whisperx.load_model("large-v2")
# 2. Process audio
audio = load_audio("meeting.wav")
diar_result = diar_pipeline(audio) # Pyannote speaker turns
asr_result = whisperx.align(audio, asr_model, metadata)
# 3. Fuse results
segments = assign_speakers_from_diarization(
asr_result["segments"],
diar_result
)
final_segments = list(resplit_segments_by_diarization(
segments,
diar_result
))
# 4. Consume downstream
for seg in final_segments:
print(f"[{seg['speaker']}] {seg['start']:.2f}-{seg['end']:.2f}: {seg['text']}")
Summary
- Pyannote (
pyannote/speaker-diarization-3.1) provides speaker boundary detection via lazily-loaded pipeline inmodel_manager.py - WhisperX delivers high-accuracy transcription with word-level timestamps through
align()in the ASR backend - Fusion logic in
segmentation.pyuses time-overlap weighting for speaker assignment and re-splitting for clean segment boundaries - Output integration feeds speaker cloning, multi-format exports, and API delivery with unified
{speaker, start, end, text}structures - Test coverage in
tests/test_assign_speakers_from_diarization.pyandtests/test_resplit_speaker.pyvalidates edge cases
Frequently Asked Questions
What version of Pyannote does VoiceStudio use?
VoiceStudio specifically uses pyannote/speaker-diarization-3.1, the current state-of-the-art release from Hugging Face Hub. The get_diarization_pipeline() function in backend/services/model_manager.py pins this model identifier and handles authentication for the gated repository.
Why use WhisperX instead of standard OpenAI Whisper?
WhisperX adds forced phonetic alignment that produces sub-second word-level timestamps. Standard Whisper only provides segment-level timing (~10-30 second chunks), which is too coarse for accurate speaker assignment when diarization boundaries fall mid-segment. The whisperx.align() call in backend/services/asr_backend.py delivers the precision needed for overlap-based fusion.
How does VoiceStudio handle overlapping speech?
The assign_speakers_from_diarization() function uses weighted temporal overlap rather than simple binary intersection. When multiple speakers overlap a transcript segment, the speaker with the greatest cumulative overlap time wins assignment. For complete overlap scenarios, the dominant speaker is annotated; advanced multi-label assignment would require extending the segmentation logic.
Can the diarization pipeline run on CPU only?
Yes, though with significant latency tradeoffs. The get_diarization_pipeline() function respects PyTorch device configuration—CUDA is preferred when available, but CPU fallback is supported. For production deployments, GPU acceleration is strongly recommended; the Pyannote diarization-3.1 model contains speaker embedding networks that benefit substantially from parallel computation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →