How the Speech-to-Speech VAD Pipeline Handles Turn-Taking and Interruption Detection with Silero VAD

The VAD pipeline uses a state-machine-driven VADHandler that wraps Silero VAD to detect speech boundaries, validate minimum active speech durations, and dynamically reopen turns when pauses fall within configurable thresholds, enabling seamless interruption detection.

The huggingface/speech-to-speech repository implements a real-time voice activity detection (VAD) subsystem built around Silero VAD (loaded from the snakers4/silero-vad hub). By combining frame-level probability scores from the Silero model with sophisticated turn-management logic, the pipeline achieves robust turn-taking and interruption detection suitable for conversational AI assistants.

Loading and Initializing the Silero VAD Model

The pipeline initializes the VAD subsystem in VADHandler.setup, which fetches the pre-trained model via torch.hub.load and instantiates a streaming-compatible VADIterator.


# From src/speech_to_speech/VAD/vad_handler.py

def setup(self, should_listen, thresh=0.5, sample_rate=16000, 
          min_silence_ms=64, min_speech_ms=384, speech_pad_ms=30, ...):
    self.model, utils = torch.hub.load(repo_or_dir='snakers4/silero-vad',
                                       model='silero_vad',
                                       force_reload=False)
    self.iterator = VADIterator(self.model, 
                                threshold=thresh,
                                sampling_rate=sample_rate,
                                min_silence_duration_ms=min_silence_ms,
                                speech_pad_ms=speech_pad_ms)

The VADIterator class (located in src/speech_to_speech/VAD/vad_iterator.py) acts as a thin wrapper that converts the Silero model into a streaming iterator, managing pre-speech buffering and end-of-utterance logic.

Detecting Speech Boundaries in Streaming Audio

Audio flows through the pipeline as raw PCM chunks converted to torch.from_numpy tensors. The VADIterator.__call__ method processes each chunk and maintains a triggered flag indicating whether the VAD has entered a speech segment.

While speech is ongoing, the iterator returns None. When an utterance ends (based on silence duration exceeding min_silence_ms), it yields a list of buffered audio tensors. This design allows the VADHandler to distinguish between transient noise and actual speech boundaries without latency-heavy post-processing.

Turn-Taking Logic and Turn Validation

Starting New Turns with Active Speech Validation

To prevent false triggers from background noise, the handler validates speech duration before declaring a new turn. In VADHandler.process, when iterator.triggered becomes true, the code checks _active_speech_min_ms to ensure sufficient speech duration has accumulated.

Only after this validation does _ensure_turn_for_speech_start invoke _start_new_turn, which assigns a monotonic turn_{n} identifier and revision number. This gatekeeping ensures that brief sounds below the configurable noise floor do not create spurious turn boundaries.

Turn Bookkeeping and Speculative Tracking

The _start_new_turn method (lines 76-81 in vad_handler.py) initializes turn state and, when speculative-turn tracking is enabled, registers the turn with a SpeculativeTurnTracker. This tracker observes turn lifecycle events and manages metadata required for the interruption detection heuristics described below.

Interruption Detection and Turn Reopening

Reopening Turns After Short Pauses

The pipeline detects interruptions by distinguishing between final turn endings and temporary pauses. If speech ceases inside an active turn, _should_reopen_current_turn evaluates whether the pause duration is within speculative_reopen_ms (or the longer unanswered_reopen_ms when the assistant has not yet replied).

If conditions are met, _reopen_current_turn reactivates the existing turn rather than creating a fresh one. This prevents the system from treating brief user hesitations as the end of a conversational turn, allowing the assistant to be interrupted naturally without losing context.

Pending Reopen Candidates for Smooth Transitions

When a speech start is detected during a potential reopen window, _begin_pending_reopen_if_needed records a candidate reopen state. If the turn-reopen conditions still hold after validation, _confirm_pending_reopen finalizes the decision. This mechanism handles edge cases where the user interrupts shortly after the assistant begins speaking, ensuring smooth transitions without audio artifacts or lost utterances.

Finalizing Utterances and Emitting Events

Once the iterator returns a non-empty buffer indicating speech end, the handler finalizes the utterance. In _process_realtime and _process_normal (lines 69-120), the pipeline optionally merges short "held" segments, applies audio enhancement, and emits structured events:

  • SpeechStartedEvent: Dispatched after turn validation, containing turn_id and audio_start_ms
  • SpeechStoppedEvent: Dispatched with turn_id and audio_end_ms when the utterance concludes
  • VADAudio: Message objects carrying audio chunks with mode "final" for downstream STT processing

These events (defined in src/speech_to_speech/pipeline/events.py) carry turn metadata that downstream components use to correlate transcriptions with specific conversational turns.

Runtime Configuration and Dynamic Thresholds

The handler supports dynamic adaptation through _apply_runtime_turn_detection (lines 45-62), which respects per-session RuntimeConfig updates. This allows adjustment of threshold and silence_duration_ms on-the-fly, enabling the system to adapt to noisy environments or user-requested sensitivity changes without restarting the pipeline.

Practical Implementation Example

Below is a minimal implementation showing how to wire the VAD pipeline into a real-time audio loop:

import threading
import queue
import numpy as np
import torch

from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.pipeline.messages import VADAudio
from speech_to_speech.pipeline.events import SpeechStartedEvent, SpeechStoppedEvent

# Initialize the handler

listen_event = threading.Event()
listen_event.set()
text_q = queue.Queue()

vad = VADHandler()
vad.setup(
    should_listen=listen_event,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    min_speech_ms=384,
    speech_pad_ms=30,
    enable_realtime_transcription=True,
    text_output_queue=text_q,
)

# Feed raw PCM chunks (e.g., from microphone callback)

def audio_callback(pcm_bytes: bytes):
    for out in vad.process(pcm_bytes):
        if isinstance(out, VADAudio):
            handle_audio(out)

# Consume turn-level events

def event_worker():
    while True:
        ev = text_q.get()
        if isinstance(ev, SpeechStartedEvent):
            print(f"Turn {ev.turn_id} started at {ev.audio_start_ms} ms")
        elif isinstance(ev, SpeechStoppedEvent):
            print(f"Turn {ev.turn_id} ended at {ev.audio_end_ms} ms")

threading.Thread(target=event_worker, daemon=True).start()

The handler automatically manages turn IDs, detects interruptions via the reopen logic, and emits SpeechStartedEvent and SpeechStoppedEvent with correct identifiers for downstream processing.

Summary

  • Silero VAD Integration: The pipeline loads the snakers4/silero-vad model via torch.hub.load and wraps it in VADIterator for streaming compatibility.
  • Turn Validation: New turns require validation against _active_speech_min_ms to filter out transient noise, implemented in _ensure_turn_for_speech_start.
  • Interruption Detection: Short pauses within speculative_reopen_ms or unanswered_reopen_ms trigger _reopen_current_turn rather than finalizing the turn.
  • Event Architecture: The system emits SpeechStartedEvent and SpeechStoppedEvent (defined in src/speech_to_speech/pipeline/events.py) to coordinate downstream STT and LLM components.
  • Dynamic Configuration: Runtime parameters can be adjusted via _apply_runtime_turn_detection without pipeline restarts.

Frequently Asked Questions

What is Silero VAD and why is it used in this pipeline?

Silero VAD is a lightweight, pre-trained voice activity detection model available via torch.hub. According to the huggingface/speech-to-speech source code, it provides high-accuracy frame-level speech probabilities with low computational overhead, making it ideal for real-time streaming applications where latency and CPU usage are critical constraints.

How does the pipeline distinguish between a pause and a turn end?

The VADHandler evaluates pause duration against configurable thresholds (speculative_reopen_ms for general cases, unanswered_reopen_ms when awaiting assistant response). If the silence duration exceeds these thresholds, _should_reopen_current_turn returns false and the turn finalizes; otherwise, the turn enters a pending reopen state via _begin_pending_reopen_if_needed.

What triggers an interruption detection versus a new turn creation?

An interruption is detected when speech resumes during the reopen window of an existing active turn, causing _reopen_current_turn to extend the current turn_id rather than invoking _start_new_turn. A new turn is created only when speech begins after a complete silence exceeding the reopen thresholds, or when no previous turn is active.

Can the VAD sensitivity be adjusted during a conversation?

Yes. The _apply_runtime_turn_detection method (lines 45-62 in vad_handler.py) accepts RuntimeConfig updates that modify threshold and silence_duration_ms dynamically. This allows applications to adjust sensitivity in response to environmental noise changes or user preferences without restarting the audio processing thread.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →