# How Silero VAD Determines Speech Boundaries and Emits speech_started / speech_stopped Events

> Learn how Silero VAD detects speech boundaries by comparing probabilities to a threshold and emits speech_started and speech_stopped events through its VADHandler.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-10

---

**Silero VAD determines speech boundaries by comparing per-chunk speech probabilities against a configurable threshold, using a hysteresis mechanism (threshold - 0.15) for silence detection, while the `VADHandler` converts these internal state changes into public `SpeechStartedEvent` and `SpeechStoppedEvent` objects.**

The huggingface/speech-to-speech repository implements a two-layer Voice Activity Detection pipeline that transforms continuous audio streams into discrete utterances. This architecture separates low-level model inference from high-level event orchestration, enabling precise boundary detection with OpenAI-Realtime-compatible event semantics.

## Architecture Overview: Two-Layer Detection

The repository separates concerns into a low-level detector and a high-level event manager:

| Layer | File | Responsibility |
|-------|------|----------------|
| **Silero VAD + VADIterator** | [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) | Converts audio chunks to speech probabilities, applies thresholding and hysteresis, returns completed utterance buffers |
| **VADHandler** | [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) | Orchestrates the iterator, manages stateful turn logic, and emits `SpeechStartedEvent` / `SpeechStoppedEvent` |

## Layer 1: Probability-Based Boundary Detection in VADIterator

The `VADIterator` class in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) wraps the Silero VAD model and implements the core boundary detection algorithm.

### Model Initialization and Configuration

During setup in `VADHandler.setup` (lines 95-107 of [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)), the Silero VAD loads via `torch.hub.load` and initializes the iterator with hyperparameters:

```python
self.model, _ = torch.hub.load(
    "snakers4/silero-vad",
    "silero_vad",
    trust_repo=True,
    skip_validation=True,
)
self.iterator = VADIterator(
    self.model,
    threshold=thresh,
    sampling_rate=sample_rate,
    min_silence_duration_ms=min_silence_ms,
    speech_pad_ms=speech_pad_ms,
)

```

### Speech Start Detection

Each incoming audio chunk converts to float32 and passes through `VADIterator.__call__`. The model returns a speech probability (line 29 of [`vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_iterator.py)):

```python
speech_prob = self.model(x, self.sampling_rate).item()

```

When `speech_prob >= threshold` and the iterator is not already triggered, the system marks the speech start:

```python
if (speech_prob >= self.threshold) and not self.triggered:
    self.triggered = True
    self.prefix_buffer = list(self._pre_speech_buffer)
    # Begin buffering current and subsequent chunks

```

The iterator maintains a `_pre_speech_buffer` controlled by `speech_pad_ms` to preserve audio milliseconds before the trigger point, preventing initial phoneme truncation.

### Silence Detection and Speech End

While `triggered` is active, the iterator monitors for silence using a hysteresis threshold of `threshold - 0.15`. When probability drops below this value (lines 53-58), the iterator records a temporary end timestamp:

```python
if speech_prob < self.threshold - 0.15:
    if not self.temp_end:
        self.temp_end = self.current_sample
    if self.current_sample - self.temp_end < self.min_silence_samples:
        return None
    # Silence duration exceeded min_silence_duration_ms → speech ended

```

If the silence persists longer than `min_silence_duration_ms`, the iterator:
1. Clears `self.triggered`
2. Returns the accumulated `speech_buffer` as a list of tensors
3. Resets internal buffers for the next utterance

## Layer 2: Event Emission in VADHandler

The `VADHandler` class in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) bridges the iterator's internal state with the OpenAI-Realtime-compatible event system.

### Audio Preprocessing

The `process` method receives raw PCM bytes and converts them (lines 103-107):

```python
audio_int16 = np.frombuffer(audio_chunk, dtype=np.int16)
audio_float32 = int2float(audio_int16)
vad_output = self.iterator(torch.from_numpy(audio_float32))

```

### Emitting SpeechStartedEvent

The handler tracks `self.iterator.triggered` to detect state transitions. When speech begins and the active duration exceeds `self._active_speech_min_ms`, it emits a `SpeechStartedEvent` (lines 111-124):

```python
is_triggered_now = self.iterator.triggered
if is_triggered_now and not self._speech_started_emitted:
    # Calculate timestamps based on _audio_ms and buffer durations

    self.text_output_queue.put(
        SpeechStartedEvent(
            audio_start_ms=effective_start_ms,
            turn_id=turn_id,
            turn_revision=turn_revision,
            reopened=reopened,
        )
    )

```

The handler calculates precise timestamps using:
- `self._audio_ms`: Total milliseconds received
- `self._speech_buffer_duration_ms()`: Buffered audio length
- `self._current_active_speech_duration_ms()`: Duration above threshold

### Emitting SpeechStoppedEvent

When `VADIterator` returns a non-empty list (`vad_output is not None`), the speech segment completes. The handler constructs the final audio array and emits `SpeechStoppedEvent` (lines 220-221 and subsequent):

```python
if vad_output is not None:
    array = torch.cat(vad_output).cpu().numpy()
    end_ms = self._audio_ms
    self.text_output_queue.put(
        SpeechStoppedEvent(
            audio_end_ms=end_ms,
            turn_id=turn_id,
            turn_revision=turn_revision,
        )
    )
    self._speech_started_emitted = False  # Reset for next utterance

```

### Real-Time Progressive Mode

When `enable_realtime_transcription=True`, the handler additionally emits `VADAudio` objects with `mode="progressive"` while speech is ongoing, enabling streaming transcription before the final `SpeechStoppedEvent`.

## Practical Implementation Examples

### Basic Synchronous Usage

```python
from speech_to_speech.VAD.vad_handler import VADHandler
from queue import Queue
from threading import Event

should_listen = Event()
should_listen.set()
output_q = Queue()

handler = VADHandler()
handler.setup(
    should_listen,
    thresh=0.5,
    sample_rate=16000,
    min_silence_ms=300,
    speech_pad_ms=30,
    text_output_queue=output_q,
)

# Process raw PCM stream

with open("sample.raw", "rb") as f:
    while chunk := f.read(640):  # 20ms @ 16kHz

        for event in handler.process(chunk):
            print(event)  # SpeechStartedEvent, SpeechStoppedEvent, or VADAudio

```

### Real-Time Configuration

```python
handler.setup(
    should_listen,
    thresh=0.6,
    sample_rate=16000,
    min_silence_ms=64,
    speech_pad_ms=30,
    enable_realtime_transcription=True,
    realtime_processing_pause=0.5,
    text_output_queue=output_q,
)

```

This configuration emits progressive audio chunks during active speech, suitable for low-latency streaming pipelines.

## Summary

- **Silero VAD** generates per-chunk speech probabilities via `torch.hub.load("snakers4/silero-vad", "silero_vad")`.
- **VADIterator** applies thresholding (start: `prob >= threshold`, end: `prob < threshold - 0.15`) with configurable `min_silence_duration_ms` to prevent false positives.
- **Pre-speech buffering** (`speech_pad_ms`) captures audio milliseconds before the trigger to avoid cutting initial phonemes.
- **VADHandler** translates internal iterator states into public `SpeechStartedEvent` and `SpeechStoppedEvent` objects, supporting OpenAI-Realtime semantics including turn reopening.
- **Real-time mode** enables progressive audio emission before speech completion, optimizing latency in streaming applications.

## Frequently Asked Questions

### What hysteresis threshold does Silero VAD use for silence detection?

The `VADIterator` uses a hysteresis of **0.15** below the configured threshold. Speech starts when probability exceeds `threshold`, but only ends when probability drops below `threshold - 0.15` for longer than `min_silence_duration_ms`. This prevents rapid toggling during brief pauses.

### How does the system prevent losing audio at the beginning of speech?

The iterator maintains a `_pre_speech_buffer` limited by `speech_pad_ms` (default 30ms). When speech triggers, this buffer prepends to the active speech buffer, ensuring the initial milliseconds of phonemes are preserved in the final utterance.

### What is the difference between VADIterator and VADHandler?

`VADIterator` ([`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py)) handles low-level model inference and binary state management (triggered/not triggered). `VADHandler` ([`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py)) provides high-level orchestration, timestamp calculation, turn management, and event emission compatible with the OpenAI Realtime API specification.

### Can Silero VAD handle different sampling rates?

Yes. The `VADIterator` accepts a `sampling_rate` parameter (typically 16000 or 8000 Hz) and passes it to the model during inference: `self.model(x, self.sampling_rate).item()`. The handler automatically configures sample rate conversion if the input stream differs.