How Silero VAD Determines Speech Boundaries and Emits speech_started / speech_stopped Events

Silero VAD determines speech boundaries by comparing per-chunk speech probabilities against a configurable threshold, using a hysteresis mechanism (threshold - 0.15) for silence detection, while the VADHandler converts these internal state changes into public SpeechStartedEvent and SpeechStoppedEvent objects.

The huggingface/speech-to-speech repository implements a two-layer Voice Activity Detection pipeline that transforms continuous audio streams into discrete utterances. This architecture separates low-level model inference from high-level event orchestration, enabling precise boundary detection with OpenAI-Realtime-compatible event semantics.

Architecture Overview: Two-Layer Detection

The repository separates concerns into a low-level detector and a high-level event manager:

Layer File Responsibility
Silero VAD + VADIterator src/speech_to_speech/VAD/vad_iterator.py Converts audio chunks to speech probabilities, applies thresholding and hysteresis, returns completed utterance buffers
VADHandler src/speech_to_speech/VAD/vad_handler.py Orchestrates the iterator, manages stateful turn logic, and emits SpeechStartedEvent / SpeechStoppedEvent

Layer 1: Probability-Based Boundary Detection in VADIterator

The VADIterator class in src/speech_to_speech/VAD/vad_iterator.py wraps the Silero VAD model and implements the core boundary detection algorithm.

Model Initialization and Configuration

During setup in VADHandler.setup (lines 95-107 of vad_handler.py), the Silero VAD loads via torch.hub.load and initializes the iterator with hyperparameters:

self.model, _ = torch.hub.load(
    "snakers4/silero-vad",
    "silero_vad",
    trust_repo=True,
    skip_validation=True,
)
self.iterator = VADIterator(
    self.model,
    threshold=thresh,
    sampling_rate=sample_rate,
    min_silence_duration_ms=min_silence_ms,
    speech_pad_ms=speech_pad_ms,
)

Speech Start Detection

Each incoming audio chunk converts to float32 and passes through VADIterator.__call__. The model returns a speech probability (line 29 of vad_iterator.py):

speech_prob = self.model(x, self.sampling_rate).item()

When speech_prob >= threshold and the iterator is not already triggered, the system marks the speech start:

if (speech_prob >= self.threshold) and not self.triggered:
    self.triggered = True
    self.prefix_buffer = list(self._pre_speech_buffer)
    # Begin buffering current and subsequent chunks

The iterator maintains a _pre_speech_buffer controlled by speech_pad_ms to preserve audio milliseconds before the trigger point, preventing initial phoneme truncation.

Silence Detection and Speech End

While triggered is active, the iterator monitors for silence using a hysteresis threshold of threshold - 0.15. When probability drops below this value (lines 53-58), the iterator records a temporary end timestamp:

if speech_prob < self.threshold - 0.15:
    if not self.temp_end:
        self.temp_end = self.current_sample
    if self.current_sample - self.temp_end < self.min_silence_samples:
        return None
    # Silence duration exceeded min_silence_duration_ms → speech ended

If the silence persists longer than min_silence_duration_ms, the iterator:

  1. Clears self.triggered
  2. Returns the accumulated speech_buffer as a list of tensors
  3. Resets internal buffers for the next utterance

Layer 2: Event Emission in VADHandler

The VADHandler class in src/speech_to_speech/VAD/vad_handler.py bridges the iterator's internal state with the OpenAI-Realtime-compatible event system.

Audio Preprocessing

The process method receives raw PCM bytes and converts them (lines 103-107):

audio_int16 = np.frombuffer(audio_chunk, dtype=np.int16)
audio_float32 = int2float(audio_int16)
vad_output = self.iterator(torch.from_numpy(audio_float32))

Emitting SpeechStartedEvent

The handler tracks self.iterator.triggered to detect state transitions. When speech begins and the active duration exceeds self._active_speech_min_ms, it emits a SpeechStartedEvent (lines 111-124):

is_triggered_now = self.iterator.triggered
if is_triggered_now and not self._speech_started_emitted:
    # Calculate timestamps based on _audio_ms and buffer durations

    self.text_output_queue.put(
        SpeechStartedEvent(
            audio_start_ms=effective_start_ms,
            turn_id=turn_id,
            turn_revision=turn_revision,
            reopened=reopened,
        )
    )

The handler calculates precise timestamps using:

  • self._audio_ms: Total milliseconds received
  • self._speech_buffer_duration_ms(): Buffered audio length
  • self._current_active_speech_duration_ms(): Duration above threshold

Emitting SpeechStoppedEvent

When VADIterator returns a non-empty list (vad_output is not None), the speech segment completes. The handler constructs the final audio array and emits SpeechStoppedEvent (lines 220-221 and subsequent):

if vad_output is not None:
    array = torch.cat(vad_output).cpu().numpy()
    end_ms = self._audio_ms
    self.text_output_queue.put(
        SpeechStoppedEvent(
            audio_end_ms=end_ms,
            turn_id=turn_id,
            turn_revision=turn_revision,
        )
    )
    self._speech_started_emitted = False  # Reset for next utterance

Real-Time Progressive Mode

When enable_realtime_transcription=True, the handler additionally emits VADAudio objects with mode="progressive" while speech is ongoing, enabling streaming transcription before the final SpeechStoppedEvent.

Practical Implementation Examples

Basic Synchronous Usage

from speech_to_speech.VAD.vad_handler import VADHandler
from queue import Queue
from threading import Event

should_listen = Event()
should_listen.set()
output_q = Queue()

handler = VADHandler()
handler.setup(
    should_listen,
    thresh=0.5,
    sample_rate=16000,
    min_silence_ms=300,
    speech_pad_ms=30,
    text_output_queue=output_q,
)

# Process raw PCM stream

with open("sample.raw", "rb") as f:
    while chunk := f.read(640):  # 20ms @ 16kHz

        for event in handler.process(chunk):
            print(event)  # SpeechStartedEvent, SpeechStoppedEvent, or VADAudio

Real-Time Configuration

handler.setup(
    should_listen,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    speech_pad_ms=30,
    enable_realtime_transcription=True,
    realtime_processing_pause=0.5,
    text_output_queue=output_q,
)

This configuration emits progressive audio chunks during active speech, suitable for low-latency streaming pipelines.

Summary

  • Silero VAD generates per-chunk speech probabilities via torch.hub.load("snakers4/silero-vad", "silero_vad").
  • VADIterator applies thresholding (start: prob >= threshold, end: prob < threshold - 0.15) with configurable min_silence_duration_ms to prevent false positives.
  • Pre-speech buffering (speech_pad_ms) captures audio milliseconds before the trigger to avoid cutting initial phonemes.
  • VADHandler translates internal iterator states into public SpeechStartedEvent and SpeechStoppedEvent objects, supporting OpenAI-Realtime semantics including turn reopening.
  • Real-time mode enables progressive audio emission before speech completion, optimizing latency in streaming applications.

Frequently Asked Questions

What hysteresis threshold does Silero VAD use for silence detection?

The VADIterator uses a hysteresis of 0.15 below the configured threshold. Speech starts when probability exceeds threshold, but only ends when probability drops below threshold - 0.15 for longer than min_silence_duration_ms. This prevents rapid toggling during brief pauses.

How does the system prevent losing audio at the beginning of speech?

The iterator maintains a _pre_speech_buffer limited by speech_pad_ms (default 30ms). When speech triggers, this buffer prepends to the active speech buffer, ensuring the initial milliseconds of phonemes are preserved in the final utterance.

What is the difference between VADIterator and VADHandler?

VADIterator (src/speech_to_speech/VAD/vad_iterator.py) handles low-level model inference and binary state management (triggered/not triggered). VADHandler (src/speech_to_speech/VAD/vad_handler.py) provides high-level orchestration, timestamp calculation, turn management, and event emission compatible with the OpenAI Realtime API specification.

Can Silero VAD handle different sampling rates?

Yes. The VADIterator accepts a sampling_rate parameter (typically 16000 or 8000 Hz) and passes it to the model during inference: self.model(x, self.sampling_rate).item(). The handler automatically configures sample rate conversion if the input stream differs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →