Configuring Silero VAD Parameters for Different Acoustic Environments in Hugging Face Speech-to-Speech

The speech-to-speech library exposes Silero Voice Activity Detection (VAD) parameters—threshold, min_silence_duration_ms, speech_pad_ms, and sample_rate—through the VADIterator and VADHandler classes, allowing real-time acoustic adaptation without pipeline restarts.

This guide covers how to tune Silero VAD settings in the huggingface/speech-to-speech repository to optimize speech segmentation for quiet studios, noisy cafés, or reverberant spaces. The implementation centers on two core components that wrap the Silero model and integrate with the OpenAI Realtime API runtime configuration.

Core VAD Components

The VAD system relies on a two-layer architecture that separates streaming logic from pipeline management.

VADIterator

The VADIterator class in src/speech_to_speech/VAD/vad_iterator.py implements the streaming VAD logic. It manages the internal state of speech detection and exposes the core parameters that control sensitivity and timing. This class validates sample rates (supporting only 8000 Hz or 16000 Hz at line 49) and converts timestamps between milliseconds and samples based on the selected rate.

VADHandler

The VADHandler class in src/speech_to_speech/VAD/vad_handler.py serves as the pipeline bridge. It loads the Silero model via torch.hub.load, instantiates a VADIterator with user-provided arguments, and adapts VAD settings at runtime through the RuntimeConfig. The handler’s _apply_runtime_turn_detection method (lines 45–71) enables dynamic updates to threshold and min_silence_samples on-the-fly.

Key Silero VAD Parameters for Acoustic Tuning

Four parameters control how aggressively the VAD segments speech and how it handles edge cases in different environments.

Threshold

The threshold parameter sets the speech probability cutoff. Lower values make the VAD more sensitive (detecting quieter speech) but increase false positives, while higher values reduce sensitivity for loud environments.

  • Quiet room: 0.3–0.5
  • Noisy café or office: 0.6–0.7

Minimum Silence Duration

The min_silence_duration_ms (exposed as min_silence_ms in VADHandler) defines the silence length required to end a speech segment. Larger values prevent premature chopping in reverberant spaces or with choppy audio.

  • Short utterances: 300 ms
  • Conversational speech: 500–800 ms

Speech Padding

The speech_pad_ms parameter controls the amount of audio pre-roll retained before the VAD trigger. Increasing this value captures leading consonants that might otherwise be clipped.

  • Default: 30 ms
  • Clipped consonants: 50–100 ms

Sample Rate

The sample_rate must be either 8000 Hz or 16000 Hz. Use 16000 Hz for high-quality streams; 8000 Hz suits bandwidth-constrained deployments.

Configuration Examples

Static Configuration for Noisy Environments

Instantiate VADHandler with custom parameters optimized for loud acoustic conditions.

from speech_to_speech.VAD.vad_handler import VADHandler
from threading import Event

listen_event = Event()
listen_event.set()

# Configure for noisy café environment

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    thresh=0.65,                 # Higher threshold reduces false triggers

    sample_rate=16000,
    min_silence_ms=600,          # Longer pause required to close segment

    speech_pad_ms=80,            # Preserve more leading audio

    audio_enhancement=False,
    enable_realtime_transcription=True,
)

The handler now processes raw 16-bit PCM chunks via its process method, applying these conservative settings to filter background chatter.

Dynamic Runtime Updates via RuntimeConfig

Adjust VAD behavior without restarting the pipeline by passing updated RuntimeConfig objects.

from speech_to_speech.api.openai_realtime.runtime_config import RuntimeConfig

# Update configuration mid-session

runtime_cfg = RuntimeConfig(session=...)

# Client sends new turn detection settings

runtime_cfg.session.audio.input.turn_detection = {
    "threshold": 0.55,
    "silence_duration_ms": 450,
}

# Apply changes immediately

handler.process((audio_chunk_bytes, runtime_cfg))

When VADHandler.process detects a RuntimeConfig, it triggers _apply_runtime_turn_detection to update iterator.threshold and iterator.min_silence_samples dynamically.

Low-Level VADIterator Usage

For direct control without the handler wrapper, use VADIterator directly with custom parameters.

import torch
from speech_to_speech.VAD.vad_iterator import VADIterator

# Load Silero model

model, _ = torch.hub.load(
    "snakers4/silero-vad",
    "silero_vad",
    trust_repo=True,
    skip_validation=True,
)

# Configure for sensitive detection

iterator = VADIterator(
    model,
    threshold=0.45,
    sampling_rate=16000,
    min_silence_duration_ms=400,
    speech_pad_ms=50,
)

# Stream audio chunks as torch tensors

for chunk in audio_stream:
    speech = iterator(torch.from_numpy(chunk))
    if speech is not None:
        process_utterance(speech)

Summary

  • Architecture: VADIterator handles streaming logic in src/speech_to_speech/VAD/vad_iterator.py, while VADHandler manages integration and runtime updates in src/speech_to_speech/VAD/vad_handler.py.
  • Threshold: Lower values (0.3–0.5) suit quiet rooms; higher values (0.6–0.7) filter noisy environments.
  • Timing: Adjust min_silence_duration_ms to prevent premature segment ending in reverberant spaces.
  • Padding: Increase speech_pad_ms to 50–100 ms when initial consonants are clipped.
  • Runtime Flexibility: Use RuntimeConfig and _apply_runtime_turn_detection to update parameters without restarting the pipeline.

Frequently Asked Questions

How do I prevent the VAD from cutting off words in a reverberant room?

Increase the min_silence_duration_ms parameter to 600–800 ms in VADHandler (or min_silence_duration_ms in VADIterator). This extends the required silence period before closing a speech segment, accommodating lingering echoes without splitting continuous utterances.

Can I change VAD sensitivity without restarting the speech-to-speech pipeline?

Yes. Pass a new RuntimeConfig object containing updated turn_detection settings to the VADHandler.process method. The handler’s _apply_runtime_turn_detection method dynamically updates iterator.threshold and silence durations on-the-fly.

What is the difference between min_silence_ms and min_silence_duration_ms?

min_silence_duration_ms is the parameter name used in the VADIterator class, while min_silence_ms is the argument name exposed in VADHandler.setup(). Both control the same underlying behavior: the minimum length of silence required to trigger the end of a speech segment.

Which sample rate should I use for telephone-quality audio?

Use 8000 Hz for bandwidth-constrained streams or telephone-quality audio. The VADIterator validates this rate at line 49 of vad_iterator.py. For high-fidelity applications, use 16000 Hz, which provides better temporal resolution for the Silero model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →