# How the Speech-to-Speech VAD Pipeline Handles Turn-Taking and Interruption Detection with Silero VAD

> Discover how the Silero VAD pipeline expertly manages turn-taking and interruption detection. Learn about state-machine VAD handling for seamless speech boundary detection and dynamic turn reopening.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-31

---

**The VAD pipeline uses a state-machine-driven `VADHandler` that wraps Silero VAD to detect speech boundaries, validate minimum active speech durations, and dynamically reopen turns when pauses fall within configurable thresholds, enabling seamless interruption detection.**

The `huggingface/speech-to-speech` repository implements a real-time voice activity detection (VAD) subsystem built around **Silero VAD** (loaded from the `snakers4/silero-vad` hub). By combining frame-level probability scores from the Silero model with sophisticated turn-management logic, the pipeline achieves robust turn-taking and interruption detection suitable for conversational AI assistants.

## Loading and Initializing the Silero VAD Model

The pipeline initializes the VAD subsystem in `VADHandler.setup`, which fetches the pre-trained model via `torch.hub.load` and instantiates a streaming-compatible `VADIterator`.

```python

# From src/speech_to_speech/VAD/vad_handler.py

def setup(self, should_listen, thresh=0.5, sample_rate=16000, 
          min_silence_ms=64, min_speech_ms=384, speech_pad_ms=30, ...):
    self.model, utils = torch.hub.load(repo_or_dir='snakers4/silero-vad',
                                       model='silero_vad',
                                       force_reload=False)
    self.iterator = VADIterator(self.model, 
                                threshold=thresh,
                                sampling_rate=sample_rate,
                                min_silence_duration_ms=min_silence_ms,
                                speech_pad_ms=speech_pad_ms)

```

The `VADIterator` class (located in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py)) acts as a thin wrapper that converts the Silero model into a streaming iterator, managing pre-speech buffering and end-of-utterance logic.

## Detecting Speech Boundaries in Streaming Audio

Audio flows through the pipeline as raw PCM chunks converted to `torch.from_numpy` tensors. The `VADIterator.__call__` method processes each chunk and maintains a `triggered` flag indicating whether the VAD has entered a speech segment.

While speech is ongoing, the iterator returns `None`. When an utterance ends (based on silence duration exceeding `min_silence_ms`), it yields a list of buffered audio tensors. This design allows the `VADHandler` to distinguish between transient noise and actual speech boundaries without latency-heavy post-processing.

## Turn-Taking Logic and Turn Validation

### Starting New Turns with Active Speech Validation

To prevent false triggers from background noise, the handler validates speech duration before declaring a new turn. In `VADHandler.process`, when `iterator.triggered` becomes true, the code checks `_active_speech_min_ms` to ensure sufficient speech duration has accumulated.

Only after this validation does `_ensure_turn_for_speech_start` invoke `_start_new_turn`, which assigns a monotonic `turn_{n}` identifier and revision number. This gatekeeping ensures that brief sounds below the configurable noise floor do not create spurious turn boundaries.

### Turn Bookkeeping and Speculative Tracking

The `_start_new_turn` method (lines 76-81 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)) initializes turn state and, when speculative-turn tracking is enabled, registers the turn with a `SpeculativeTurnTracker`. This tracker observes turn lifecycle events and manages metadata required for the interruption detection heuristics described below.

## Interruption Detection and Turn Reopening

### Reopening Turns After Short Pauses

The pipeline detects interruptions by distinguishing between final turn endings and temporary pauses. If speech ceases inside an active turn, `_should_reopen_current_turn` evaluates whether the pause duration is within `speculative_reopen_ms` (or the longer `unanswered_reopen_ms` when the assistant has not yet replied).

If conditions are met, `_reopen_current_turn` reactivates the existing turn rather than creating a fresh one. This prevents the system from treating brief user hesitations as the end of a conversational turn, allowing the assistant to be interrupted naturally without losing context.

### Pending Reopen Candidates for Smooth Transitions

When a speech start is detected during a potential reopen window, `_begin_pending_reopen_if_needed` records a candidate reopen state. If the turn-reopen conditions still hold after validation, `_confirm_pending_reopen` finalizes the decision. This mechanism handles edge cases where the user interrupts shortly after the assistant begins speaking, ensuring smooth transitions without audio artifacts or lost utterances.

## Finalizing Utterances and Emitting Events

Once the iterator returns a non-empty buffer indicating speech end, the handler finalizes the utterance. In `_process_realtime` and `_process_normal` (lines 69-120), the pipeline optionally merges short "held" segments, applies audio enhancement, and emits structured events:

- **`SpeechStartedEvent`**: Dispatched after turn validation, containing `turn_id` and `audio_start_ms`
- **`SpeechStoppedEvent`**: Dispatched with `turn_id` and `audio_end_ms` when the utterance concludes
- **`VADAudio`**: Message objects carrying audio chunks with mode `"final"` for downstream STT processing

These events (defined in [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py)) carry turn metadata that downstream components use to correlate transcriptions with specific conversational turns.

## Runtime Configuration and Dynamic Thresholds

The handler supports dynamic adaptation through `_apply_runtime_turn_detection` (lines 45-62), which respects per-session `RuntimeConfig` updates. This allows adjustment of `threshold` and `silence_duration_ms` on-the-fly, enabling the system to adapt to noisy environments or user-requested sensitivity changes without restarting the pipeline.

## Practical Implementation Example

Below is a minimal implementation showing how to wire the VAD pipeline into a real-time audio loop:

```python
import threading
import queue
import numpy as np
import torch

from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.pipeline.messages import VADAudio
from speech_to_speech.pipeline.events import SpeechStartedEvent, SpeechStoppedEvent

# Initialize the handler

listen_event = threading.Event()
listen_event.set()
text_q = queue.Queue()

vad = VADHandler()
vad.setup(
    should_listen=listen_event,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    min_speech_ms=384,
    speech_pad_ms=30,
    enable_realtime_transcription=True,
    text_output_queue=text_q,
)

# Feed raw PCM chunks (e.g., from microphone callback)

def audio_callback(pcm_bytes: bytes):
    for out in vad.process(pcm_bytes):
        if isinstance(out, VADAudio):
            handle_audio(out)

# Consume turn-level events

def event_worker():
    while True:
        ev = text_q.get()
        if isinstance(ev, SpeechStartedEvent):
            print(f"Turn {ev.turn_id} started at {ev.audio_start_ms} ms")
        elif isinstance(ev, SpeechStoppedEvent):
            print(f"Turn {ev.turn_id} ended at {ev.audio_end_ms} ms")

threading.Thread(target=event_worker, daemon=True).start()

```

The handler automatically manages turn IDs, detects interruptions via the reopen logic, and emits `SpeechStartedEvent` and `SpeechStoppedEvent` with correct identifiers for downstream processing.

## Summary

- **Silero VAD Integration**: The pipeline loads the `snakers4/silero-vad` model via `torch.hub.load` and wraps it in `VADIterator` for streaming compatibility.
- **Turn Validation**: New turns require validation against `_active_speech_min_ms` to filter out transient noise, implemented in `_ensure_turn_for_speech_start`.
- **Interruption Detection**: Short pauses within `speculative_reopen_ms` or `unanswered_reopen_ms` trigger `_reopen_current_turn` rather than finalizing the turn.
- **Event Architecture**: The system emits `SpeechStartedEvent` and `SpeechStoppedEvent` (defined in [`src/speech_to_speech/pipeline/events.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/events.py)) to coordinate downstream STT and LLM components.
- **Dynamic Configuration**: Runtime parameters can be adjusted via `_apply_runtime_turn_detection` without pipeline restarts.

## Frequently Asked Questions

### What is Silero VAD and why is it used in this pipeline?

**Silero VAD** is a lightweight, pre-trained voice activity detection model available via `torch.hub`. According to the `huggingface/speech-to-speech` source code, it provides high-accuracy frame-level speech probabilities with low computational overhead, making it ideal for real-time streaming applications where latency and CPU usage are critical constraints.

### How does the pipeline distinguish between a pause and a turn end?

The `VADHandler` evaluates pause duration against configurable thresholds (`speculative_reopen_ms` for general cases, `unanswered_reopen_ms` when awaiting assistant response). If the silence duration exceeds these thresholds, `_should_reopen_current_turn` returns false and the turn finalizes; otherwise, the turn enters a pending reopen state via `_begin_pending_reopen_if_needed`.

### What triggers an interruption detection versus a new turn creation?

An interruption is detected when speech resumes during the reopen window of an existing active turn, causing `_reopen_current_turn` to extend the current `turn_id` rather than invoking `_start_new_turn`. A new turn is created only when speech begins after a complete silence exceeding the reopen thresholds, or when no previous turn is active.

### Can the VAD sensitivity be adjusted during a conversation?

Yes. The `_apply_runtime_turn_detection` method (lines 45-62 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)) accepts `RuntimeConfig` updates that modify `threshold` and `silence_duration_ms` dynamically. This allows applications to adjust sensitivity in response to environmental noise changes or user preferences without restarting the audio processing thread.