Huggingface Speech-to-Speech VAD Turn-Taking Parameters: Configuring --thresh, --min_speech_ms, and --min_silence_ms

The --thresh, --min_speech_ms, and --min_silence_ms command-line arguments control Voice Activity Detection (VAD) turn-taking by setting the confidence threshold for speech detection, the minimum duration of active speech required to start a turn, and the silence interval needed to close a speech segment.

The huggingface/speech-to-speech repository implements real-time conversation flow through three critical VAD turn-taking parameters that determine when user speech begins and ends. These parameters, accessed via VADHandlerArguments and applied in VADHandler, prevent false triggers from background noise and ensure natural pause detection. Understanding how to configure --thresh, --min_speech_ms, and --min_silence_ms allows precise tuning of turn boundaries for voice-based AI interactions.

How --thresh Controls Speech Detection Confidence

The --thresh parameter sets the probability threshold for the underlying Silero VAD model, with a default value of 0.6. Stored in VADHandlerArguments.thresh, this value is assigned to VADIterator.threshold during the setup() method in src/speech_to_speech/VAD/vad_handler.py. Frames with speech probability below this threshold are classified as silence, while frames above it indicate active speech.

Higher values make the system more conservative, reducing false positives but potentially delaying detection of low-energy speech. Lower values increase sensitivity to quiet speech but may trigger on non-speech noise. The handler can update this threshold at runtime via _apply_runtime_turn_detection(), enabling adaptive sensitivity without pipeline restarts.

How --min_speech_ms Filters Short Utterances

The --min_speech_ms parameter defines the minimum length of active speech required before the system emits a SpeechStartedEvent, defaulting to 384 ms. This value, stored as VADHandlerArguments.min_speech_ms, prevents brief noises or false positives from triggering new conversation turns.

In src/speech_to_speech/VAD/vad_handler.py, the handler consults this parameter in _active_speech_min_ms and _ensure_turn_for_speech_start to verify that effective_active_speech_duration_ms >= active_speech_min_ms before confirming a turn has started. The parameter also supports continuation logic through min_speech_continuation_ms, allowing shorter follow-up utterances when a turn is already active.

How --min_silence_ms Defines Turn Boundaries

The --min_silence_ms parameter specifies the shortest silence interval that the VAD treats as a break between speech segments, with a default of 64 ms. When silence exceeds this duration, the current speech segment closes and the handler emits a SpeechStoppedEvent, signaling that a new turn may begin.

Stored in VADHandlerArguments.min_silence_ms, this value is passed to VADIterator.min_silence_duration_ms in the setup() method. Pauses shorter than this threshold keep the speech segment open, effectively merging consecutive utterances into a single continuous turn.

Source Code Implementation

The turn-taking logic is distributed across three key files:

Configuration Examples

Direct Handler Instantiation

from speech_to_speech.VAD.vad_handler import VADHandler
from threading import Event

listen_event = Event()
listen_event.set()

handler = VADHandler()
handler.setup(
    should_listen=listen_event,
    thresh=0.75,            # More conservative detection

    min_speech_ms=500,      # Require 500ms to start turn

    min_silence_ms=120,     # End turn after 120ms silence

)

Using Argument Classes

from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments

args = VADHandlerArguments(
    thresh=0.55,            # More sensitive

    min_speech_ms=300,      # Shutter utterances accepted

    min_silence_ms=80,      # Shorter pauses allowed

)

print(args)  # VADHandlerArguments(thresh=0.55, ..., min_speech_ms=300, min_silence_ms=80)

Runtime Updates

# Update parameters dynamically without restarting

handler._apply_runtime_turn_detection(new_config)

Summary

  • --thresh (default 0.6) sets the VAD model confidence threshold in VADIterator.threshold; higher values reduce false positives but may miss quiet speech.
  • --min_speech_ms (default 384 ms) prevents short noises from triggering turns by requiring minimum active speech duration before emitting SpeechStartedEvent.
  • --min_silence_ms (default 64 ms) determines the pause length needed to close a speech segment and trigger SpeechStoppedEvent, allowing a new turn to begin.
  • These parameters are defined in VADHandlerArguments, applied in VADHandler.setup(), and executed in VADIterator within src/speech_to_speech/VAD/.

Frequently Asked Questions

What happens if --min_speech_ms is set too low?

Setting --min_speech_ms below the default 384 ms increases responsiveness to short utterances but raises the risk of triggering turns on coughs, breaths, or transient background noise. Values below 100 ms may cause excessive turn-taking events and fragment the conversation flow.

How does --min_silence_ms affect perceived latency?

The --min_silence_ms parameter directly impacts how quickly the system recognizes a completed turn. Lower values (e.g., 50 ms) reduce end-of-turn latency but risk splitting single utterances with natural pauses into multiple turns, while higher values (e.g., 300 ms) ensure complete captures but introduce noticeable gaps before the system responds.

Can VAD parameters be changed without restarting the application?

Yes. According to the implementation in src/speech_to_speech/VAD/vad_handler.py, the VADHandler._apply_runtime_turn_detection() method supports dynamic updates to thresh and other turn-detection parameters during active audio processing sessions.

Where are the default values for these VAD parameters defined?

Default values are defined in src/speech_to_speech/arguments_classes/vad_arguments.py within the VADHandlerArguments dataclass: thresh=0.6, min_speech_ms=384, and min_silence_ms=64.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →