# How VoiceStudio Uses Pyannote and WhisperX for Speaker Diarization: Complete Pipeline Guide

> Discover how VoiceStudio leverages Pyannote and WhisperX for advanced speaker diarization. Learn about the pipeline integration for accurate speaker-attributed transcripts and downstream applications.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-06

---

**VoiceStudio combines a Pyannote audio diarization model with the WhisperX ASR engine in a two-stage pipeline, then merges their outputs through time-weighted overlap matching and segment re-splitting to produce speaker-attributed transcripts.** The integrated results power downstream speaker cloning, subtitle exports, and API delivery.

Speaker diarization—the task of answering "who spoke when?"—is a core capability of the open-source [VoiceStudio](https://github.com/debpalash/VoiceStudio) project. This article breaks down exactly how the repository leverages **Pyannote** for speaker boundary detection and **WhisperX** for transcription, then explains the fusion logic that unifies both streams. All technical details are derived directly from the source code implementation.

---

## Loading the Pyannote Diarization Pipeline

The diarization model is managed through lazy initialization in [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py). The function `get_diarization_pipeline()` handles the complete setup lifecycle.

```python

# backend/services/model_manager.py (lines 3465-3510)

def get_diarization_pipeline():
    """
    Lazily loads and caches the pretrained Pyannote speaker-diarization-3.1 model.
    Handles Hugging Face Hub token compatibility and safe-globals registration.
    """

```

**Key responsibilities of this loader:**

- **Model specification**: Downloads `pyannote/speaker-diarization-3.1` from Hugging Face Hub
- **Token authentication**: Manages HF Hub access tokens for gated model access
- **Safe globals registration**: Configures `torch.serialization` for secure checkpoint loading
- **Pipeline caching**: Returns a cached instance after first call to avoid repeated model loading

The pipeline object returned is a ready-to-use Pyannote `Pipeline` that accepts audio tensors and returns speaker turn annotations with start/end timestamps and speaker labels.

---

## Running WhisperX for Time-Aligned Transcription

WhisperX serves as the default **ASR backend** in [`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py). Unlike standard Whisper, WhisperX adds forced alignment for precise word-level timestamps—critical for accurate speaker assignment.

```python

# backend/services/asr_backend.py (lines 500-560)

import whisperx

# Load the ASR model

model, metadata = whisperx.load_model("large-v2", device="cuda")

# Obtain aligned transcription with word-level timestamps

result = whisperx.align(audio, model, metadata)

```

**WhisperX output format:**

| Field | Description |
|-------|-------------|
| `segments` | List of transcript segments with `start`, `end`, `text` |
| `word_segments` | Individual words with sub-second timing precision |
| `language` | Detected or specified language code |

The alignment step is what makes downstream fusion possible—without accurate timestamps, speaker assignment would fail.

---

## Merging Diarization and Transcription Results

The core integration logic lives in [`backend/services/segmentation.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/segmentation.py). Two complementary functions handle the fusion: **`assign_speakers_from_diarization()`** and **`resplit_segments_by_diarization()`**.

### Speaker Assignment via Time-Overlap Weighting

`assign_speakers_from_diarization()` (lines 558-595) replaces WhisperX's placeholder speaker IDs with actual Pyannote speaker labels:

```python

# backend/services/segmentation.py

def assign_speakers_from_diarization(asr_segments, diar_pipeline):
    """
    Assigns each transcript segment to the Pyannote speaker 
    with maximum temporal overlap.
    """
    for segment in asr_segments:
        # Calculate weighted overlap with each Pyannote speaker turn

        overlaps = compute_overlap_weights(
            segment["start"], 
            segment["end"],
            diar_pipeline(audio)  # speaker turns with (start, end, speaker)

        )
        segment["speaker"] = argmax(overlaps)  # most overlapping speaker

    return asr_segments

```

**Overlap-weighted strategy explained:**

- For each WhisperX transcript segment, compute temporal intersection with every Pyannote speaker turn
- Weight by duration of overlap, not just binary intersection
- Assign to the speaker with highest cumulative overlap score

This handles edge cases where segment boundaries don't perfectly align with speaker changes.

### Segment Re-Splitting at Speaker Boundaries

`resplit_segments_by_diarization()` (lines 785-820) performs finer-grained segmentation:

```python

# backend/services/segmentation.py

def resplit_segments_by_diarization(segments, diar_pipeline):
    """
    Re-splits transcript segments when Pyannote detects speaker changes
    mid-segment, preserving WhisperX word-level timing.
    """
    for segment in segments:
        speaker_turns = diar_pipeline.get_turns(
            segment["start"], 
            segment["end"]
        )
        if len(speaker_turns) > 1:
            # Split this segment at speaker change points

            sub_segments = split_at_boundaries(
                segment, 
                [t.end for t in speaker_turns[:-1]]
            )
            yield from sub_segments
        else:
            yield segment

```

**Why re-splitting matters:**

- WhisperX segments may span multiple speaker turns
- Clean subtitle files require speaker changes at precise timestamps
- Word-level timing from WhisperX is preserved through the split operation

---

## Downstream Integration: Where Diarized Output Flows

The unified speaker-attributed transcript (`{speaker, start, end, text}`) feeds multiple downstream systems:

### Speaker Cloning Service

[`backend/services/speaker_clone.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/speaker_clone.py) (lines 80-110) maps diarized speaker IDs to voice models:

```python

# Map detected speaker labels to clone models

speaker_id = segment["speaker"]  # e.g., "SPEAKER_01"

clone_model = speaker_registry.get_or_create(speaker_id)

```

Each unique speaker detected by Pyannote becomes a target for voice cloning operations.

### Export Formats and API Delivery

The final segments render to multiple output formats:

- **SRT/VTT**: Speaker labels prepended to subtitle text (`[SPEAKER_01] Hello world`)
- **JSON**: Structured array with full metadata for UI consumption
- **API response**: Direct passthrough to frontend for real-time display

---

## Complete Pipeline Example

```python
from services.model_manager import get_diarization_pipeline
from services.segmentation import (
    assign_speakers_from_diarization,
    resplit_segments_by_diarization
)
import whisperx

# 1. Load models (cached after first call)

diar_pipeline = get_diarization_pipeline()
asr_model, metadata = whisperx.load_model("large-v2")

# 2. Process audio

audio = load_audio("meeting.wav")
diar_result = diar_pipeline(audio)  # Pyannote speaker turns

asr_result = whisperx.align(audio, asr_model, metadata)

# 3. Fuse results

segments = assign_speakers_from_diarization(
    asr_result["segments"], 
    diar_result
)
final_segments = list(resplit_segments_by_diarization(
    segments, 
    diar_result
))

# 4. Consume downstream

for seg in final_segments:
    print(f"[{seg['speaker']}] {seg['start']:.2f}-{seg['end']:.2f}: {seg['text']}")

```

---

## Summary

- **Pyannote** (`pyannote/speaker-diarization-3.1`) provides speaker boundary detection via lazily-loaded pipeline in [`model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/model_manager.py)
- **WhisperX** delivers high-accuracy transcription with word-level timestamps through `align()` in the ASR backend
- **Fusion logic** in [`segmentation.py`](https://github.com/debpalash/VoiceStudio/blob/main/segmentation.py) uses time-overlap weighting for speaker assignment and re-splitting for clean segment boundaries
- **Output integration** feeds speaker cloning, multi-format exports, and API delivery with unified `{speaker, start, end, text}` structures
- **Test coverage** in [`tests/test_assign_speakers_from_diarization.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_assign_speakers_from_diarization.py) and [`tests/test_resplit_speaker.py`](https://github.com/debpalash/VoiceStudio/blob/main/tests/test_resplit_speaker.py) validates edge cases

---

## Frequently Asked Questions

### What version of Pyannote does VoiceStudio use?

VoiceStudio specifically uses `pyannote/speaker-diarization-3.1`, the current state-of-the-art release from Hugging Face Hub. The `get_diarization_pipeline()` function in [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py) pins this model identifier and handles authentication for the gated repository.

### Why use WhisperX instead of standard OpenAI Whisper?

WhisperX adds forced phonetic alignment that produces **sub-second word-level timestamps**. Standard Whisper only provides segment-level timing (~10-30 second chunks), which is too coarse for accurate speaker assignment when diarization boundaries fall mid-segment. The `whisperx.align()` call in [`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py) delivers the precision needed for overlap-based fusion.

### How does VoiceStudio handle overlapping speech?

The `assign_speakers_from_diarization()` function uses **weighted temporal overlap** rather than simple binary intersection. When multiple speakers overlap a transcript segment, the speaker with the greatest cumulative overlap time wins assignment. For complete overlap scenarios, the dominant speaker is annotated; advanced multi-label assignment would require extending the segmentation logic.

### Can the diarization pipeline run on CPU only?

Yes, though with significant latency tradeoffs. The `get_diarization_pipeline()` function respects PyTorch device configuration—CUDA is preferred when available, but CPU fallback is supported. For production deployments, GPU acceleration is strongly recommended; the Pyannote diarization-3.1 model contains speaker embedding networks that benefit substantially from parallel computation.