How Silero VAD Determines Speech Boundaries and Emits speech_started / speech_stopped Events
Silero VAD determines speech boundaries by comparing per-chunk speech probabilities against a configurable threshold, using a hysteresis mechanism (threshold - 0.15) for silence detection, while the VADHandler converts these internal state changes into public SpeechStartedEvent and SpeechStoppedEvent objects.
The huggingface/speech-to-speech repository implements a two-layer Voice Activity Detection pipeline that transforms continuous audio streams into discrete utterances. This architecture separates low-level model inference from high-level event orchestration, enabling precise boundary detection with OpenAI-Realtime-compatible event semantics.
Architecture Overview: Two-Layer Detection
The repository separates concerns into a low-level detector and a high-level event manager:
| Layer | File | Responsibility |
|---|---|---|
| Silero VAD + VADIterator | src/speech_to_speech/VAD/vad_iterator.py |
Converts audio chunks to speech probabilities, applies thresholding and hysteresis, returns completed utterance buffers |
| VADHandler | src/speech_to_speech/VAD/vad_handler.py |
Orchestrates the iterator, manages stateful turn logic, and emits SpeechStartedEvent / SpeechStoppedEvent |
Layer 1: Probability-Based Boundary Detection in VADIterator
The VADIterator class in src/speech_to_speech/VAD/vad_iterator.py wraps the Silero VAD model and implements the core boundary detection algorithm.
Model Initialization and Configuration
During setup in VADHandler.setup (lines 95-107 of vad_handler.py), the Silero VAD loads via torch.hub.load and initializes the iterator with hyperparameters:
self.model, _ = torch.hub.load(
"snakers4/silero-vad",
"silero_vad",
trust_repo=True,
skip_validation=True,
)
self.iterator = VADIterator(
self.model,
threshold=thresh,
sampling_rate=sample_rate,
min_silence_duration_ms=min_silence_ms,
speech_pad_ms=speech_pad_ms,
)
Speech Start Detection
Each incoming audio chunk converts to float32 and passes through VADIterator.__call__. The model returns a speech probability (line 29 of vad_iterator.py):
speech_prob = self.model(x, self.sampling_rate).item()
When speech_prob >= threshold and the iterator is not already triggered, the system marks the speech start:
if (speech_prob >= self.threshold) and not self.triggered:
self.triggered = True
self.prefix_buffer = list(self._pre_speech_buffer)
# Begin buffering current and subsequent chunks
The iterator maintains a _pre_speech_buffer controlled by speech_pad_ms to preserve audio milliseconds before the trigger point, preventing initial phoneme truncation.
Silence Detection and Speech End
While triggered is active, the iterator monitors for silence using a hysteresis threshold of threshold - 0.15. When probability drops below this value (lines 53-58), the iterator records a temporary end timestamp:
if speech_prob < self.threshold - 0.15:
if not self.temp_end:
self.temp_end = self.current_sample
if self.current_sample - self.temp_end < self.min_silence_samples:
return None
# Silence duration exceeded min_silence_duration_ms → speech ended
If the silence persists longer than min_silence_duration_ms, the iterator:
- Clears
self.triggered - Returns the accumulated
speech_bufferas a list of tensors - Resets internal buffers for the next utterance
Layer 2: Event Emission in VADHandler
The VADHandler class in src/speech_to_speech/VAD/vad_handler.py bridges the iterator's internal state with the OpenAI-Realtime-compatible event system.
Audio Preprocessing
The process method receives raw PCM bytes and converts them (lines 103-107):
audio_int16 = np.frombuffer(audio_chunk, dtype=np.int16)
audio_float32 = int2float(audio_int16)
vad_output = self.iterator(torch.from_numpy(audio_float32))
Emitting SpeechStartedEvent
The handler tracks self.iterator.triggered to detect state transitions. When speech begins and the active duration exceeds self._active_speech_min_ms, it emits a SpeechStartedEvent (lines 111-124):
is_triggered_now = self.iterator.triggered
if is_triggered_now and not self._speech_started_emitted:
# Calculate timestamps based on _audio_ms and buffer durations
self.text_output_queue.put(
SpeechStartedEvent(
audio_start_ms=effective_start_ms,
turn_id=turn_id,
turn_revision=turn_revision,
reopened=reopened,
)
)
The handler calculates precise timestamps using:
self._audio_ms: Total milliseconds receivedself._speech_buffer_duration_ms(): Buffered audio lengthself._current_active_speech_duration_ms(): Duration above threshold
Emitting SpeechStoppedEvent
When VADIterator returns a non-empty list (vad_output is not None), the speech segment completes. The handler constructs the final audio array and emits SpeechStoppedEvent (lines 220-221 and subsequent):
if vad_output is not None:
array = torch.cat(vad_output).cpu().numpy()
end_ms = self._audio_ms
self.text_output_queue.put(
SpeechStoppedEvent(
audio_end_ms=end_ms,
turn_id=turn_id,
turn_revision=turn_revision,
)
)
self._speech_started_emitted = False # Reset for next utterance
Real-Time Progressive Mode
When enable_realtime_transcription=True, the handler additionally emits VADAudio objects with mode="progressive" while speech is ongoing, enabling streaming transcription before the final SpeechStoppedEvent.
Practical Implementation Examples
Basic Synchronous Usage
from speech_to_speech.VAD.vad_handler import VADHandler
from queue import Queue
from threading import Event
should_listen = Event()
should_listen.set()
output_q = Queue()
handler = VADHandler()
handler.setup(
should_listen,
thresh=0.5,
sample_rate=16000,
min_silence_ms=300,
speech_pad_ms=30,
text_output_queue=output_q,
)
# Process raw PCM stream
with open("sample.raw", "rb") as f:
while chunk := f.read(640): # 20ms @ 16kHz
for event in handler.process(chunk):
print(event) # SpeechStartedEvent, SpeechStoppedEvent, or VADAudio
Real-Time Configuration
handler.setup(
should_listen,
thresh=0.6,
sample_rate=16000,
min_silence_ms=64,
speech_pad_ms=30,
enable_realtime_transcription=True,
realtime_processing_pause=0.5,
text_output_queue=output_q,
)
This configuration emits progressive audio chunks during active speech, suitable for low-latency streaming pipelines.
Summary
- Silero VAD generates per-chunk speech probabilities via
torch.hub.load("snakers4/silero-vad", "silero_vad"). - VADIterator applies thresholding (start:
prob >= threshold, end:prob < threshold - 0.15) with configurablemin_silence_duration_msto prevent false positives. - Pre-speech buffering (
speech_pad_ms) captures audio milliseconds before the trigger to avoid cutting initial phonemes. - VADHandler translates internal iterator states into public
SpeechStartedEventandSpeechStoppedEventobjects, supporting OpenAI-Realtime semantics including turn reopening. - Real-time mode enables progressive audio emission before speech completion, optimizing latency in streaming applications.
Frequently Asked Questions
What hysteresis threshold does Silero VAD use for silence detection?
The VADIterator uses a hysteresis of 0.15 below the configured threshold. Speech starts when probability exceeds threshold, but only ends when probability drops below threshold - 0.15 for longer than min_silence_duration_ms. This prevents rapid toggling during brief pauses.
How does the system prevent losing audio at the beginning of speech?
The iterator maintains a _pre_speech_buffer limited by speech_pad_ms (default 30ms). When speech triggers, this buffer prepends to the active speech buffer, ensuring the initial milliseconds of phonemes are preserved in the final utterance.
What is the difference between VADIterator and VADHandler?
VADIterator (src/speech_to_speech/VAD/vad_iterator.py) handles low-level model inference and binary state management (triggered/not triggered). VADHandler (src/speech_to_speech/VAD/vad_handler.py) provides high-level orchestration, timestamp calculation, turn management, and event emission compatible with the OpenAI Realtime API specification.
Can Silero VAD handle different sampling rates?
Yes. The VADIterator accepts a sampling_rate parameter (typically 16000 or 8000 Hz) and passes it to the model during inference: self.model(x, self.sampling_rate).item(). The handler automatically configures sample rate conversion if the input stream differs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →