How to Configure VAD Parameters in Hugging Face Speech-to-Speech: Thresholds, Silence Detection, and Realtime Settings
Configure VAD parameters by instantiating VADHandlerArguments with your desired thresholds and timings, then pass it to SpeechToSpeechPipeline, or update settings dynamically via the OpenAI-compatible realtime API's session.update event.
The Hugging Face speech-to-speech repository wraps the Silero VAD model inside VADHandler to detect voice activity in audio streams. All user-controllable VAD options are centralized in the VADHandlerArguments dataclass and applied through the handler's setup() method, giving you precise control over speech boundaries, silence detection, and realtime processing behavior.
Understanding VADHandlerArguments
The VADHandlerArguments dataclass in src/speech_to_speech/arguments_classes/vad_arguments.py exposes every tunable knob for voice activity detection. When you instantiate SpeechToSpeechPipeline, these arguments are forwarded to VADHandler.setup() in src/speech_to_speech/VAD/vad_handler.py, which initializes the internal VADIterator.
Speech Detection Thresholds
The thresh parameter (defined at vad_arguments.py#L6-L10) controls the confidence level required for the Silero model to consider audio as speech. It accepts float values between 0.0 and 1.0, with a default of 0.6. Higher values reduce false positives but may miss quiet speech, while lower values increase sensitivity to background noise.
Speech and Silence Duration Controls
Several parameters govern the minimum and maximum durations for speech segments and silence gaps:
min_silence_ms(default 64 ms): Minimum continuous silence that triggers a segment split. Defined atvad_arguments.py#L18-L22.min_speech_ms(default 384 ms): Minimum length for a valid speech utterance. Shorter fragments are held for merging or discarded. Defined atvad_arguments.py#L24-L28.min_speech_continuation_ms(default 192 ms): Hysteresis threshold for reopening a speech turn in realtime mode. Must be less than or equal tomin_speech_ms. Defined atvad_arguments.py#L30-L34.max_speech_ms(default ∞): Hard limit on segment duration; forces a split when exceeded. Defined atvad_arguments.py#L36-L40.speech_pad_ms(default 500 ms in arguments, overridable to 30 ms in setup): Amount of audio prepended to each segment to capture speech onsets. Defined atvad_arguments.py#L42-L46.
Realtime Processing and Enhancement Options
For streaming applications, the following settings control progressive transcription and audio preprocessing:
enable_realtime_transcription(default False): Emits audio chunks while the user is still speaking. Defined atvad_arguments.py#L54-L57.realtime_processing_pause(default 0.5 s): Base interval between progressive chunks, automatically scaled by speech length. Defined atvad_arguments.py#L58-L62.speculative_reopen_ms(default 1000 ms) andunanswered_reopen_ms(default 7000 ms): Duration that a soft-ended turn remains reopenable. Defined atvad_arguments.py#L64-L74.short_segment_merge_ms(default 0): Window allowing tiny fragments to be stitched together. Useful whenmin_silence_msis set very low. Defined atvad_arguments.py#L76-L80.audio_enhancement(default False): Enables DeepFilterNet noise reduction. Requires the optionaldfpackage. Defined atvad_arguments.py#L48-L53.
Initializing VAD Parameters at Startup
Pass a configured VADHandlerArguments instance when constructing your pipeline. The setup() method receives these values and instantiates the VADIterator with the specified threshold, sampling rate, and silence duration:
from speech_to_speech.pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
vad_args = VADHandlerArguments(
thresh=0.75, # stricter detection
min_silence_ms=100, # longer silence before split
min_speech_ms=250, # accept shorter utterances
speech_pad_ms=200, # keep only 200 ms of pre-speech context
audio_enhancement=True, # enable noise reduction (requires df)
)
pipeline = SpeechToSpeechPipeline(
vad_handler_args=vad_args,
# … other arguments like model, tokenizer, etc.
)
Inside VADHandler.setup() (lines 59–77 of vad_handler.py), these arguments populate the VADIterator:
self.iterator = VADIterator(
self.model,
threshold=thresh,
sampling_rate=sample_rate,
min_silence_duration_ms=min_silence_ms,
speech_pad_ms=speech_pad_ms,
)
Updating VAD Configuration at Runtime
When using the OpenAI-compatible realtime endpoint, clients can adjust VAD parameters without restarting the server. The handler watches for session.update events and applies changes via _apply_runtime_turn_detection (lines 45–74 of vad_handler.py).
Send a JSON message over the websocket to modify the threshold or silence duration:
# Assuming you have a websocket session `session`
session.send({
"type": "session.update",
"session": {
"audio": {
"input": {
"turn_detection": {
"type": "server_vad",
"threshold": 0.8, # raise confidence
"silence_duration_ms": 120 # lengthen silence window
}
}
}
}
})
The handler updates self.iterator.threshold and self.iterator.min_silence_samples immediately upon receiving the event.
Handling Edge Cases: Short Segments and Audio Enhancement
Merging Short Speech Segments
When min_silence_ms is set aggressively low, you may generate many tiny fragments. Enable stitching by setting short_segment_merge_ms to a positive value:
vad_args = VADHandlerArguments(
short_segment_merge_ms=150, # allow fragments within 150 ms to be merged
)
pipeline = SpeechToSpeechPipeline(vad_handler_args=vad_args)
With this setting, the handler invokes _hold_short_segment (lines 107–115) and _merge_pending_short_segment (lines 83–106) to combine fragments before discarding them.
Enabling Progressive Realtime Transcription
For low-latency streaming, enable progressive audio release:
vad_args = VADHandlerArguments(
enable_realtime_transcription=True,
realtime_processing_pause=0.3, # emit chunks every ~300 ms (scaled automatically)
)
pipeline = SpeechToSpeechPipeline(vad_handler_args=vad_args)
When enabled, the handler enters the _process_realtime branch (beginning at line 69 of vad_handler.py), yielding VADAudio objects in "progressive" mode while the user is still speaking.
Summary
- Configuration Source: All VAD defaults live in
VADHandlerArguments(src/speech_to_speech/arguments_classes/vad_arguments.py), while runtime logic resides inVADHandler(src/speech_to_speech/VAD/vad_handler.py). - Threshold Control: Adjust
thresh(0.0–1.0) to balance between false positives and missed speech. - Segment Boundaries: Use
min_silence_ms,min_speech_ms, andspeech_pad_msto define how audio is split and padded. - Realtime Updates: Modify
thresholdandsilence_duration_mson-the-fly via the OpenAI-compatible API'ssession.updateevent, processed by_apply_runtime_turn_detection. - Advanced Features: Enable
audio_enhancementfor noise reduction,short_segment_merge_msfor fragment stitching, andenable_realtime_transcriptionfor progressive streaming.
Frequently Asked Questions
What is the default VAD threshold in the speech-to-speech pipeline?
The default thresh value is 0.6, defined in src/speech_to_speech/arguments_classes/vad_arguments.py. This means the Silero model must output a probability of at least 60% for the audio to be classified as speech. You can override this at initialization or update it at runtime via the realtime API.
How do I enable realtime transcription with custom VAD timing?
Set enable_realtime_transcription=True in your VADHandlerArguments and adjust realtime_processing_pause to control the base interval between chunks. For example, setting realtime_processing_pause=0.3 yields emissions approximately every 300 milliseconds, scaled automatically by the handler based on speech length.
Can I change VAD settings without restarting the server?
Yes. When using the OpenAI-compatible realtime endpoint, send a session.update event with the turn_detection payload. The VADHandler._apply_runtime_turn_detection method (lines 45–74 of vad_handler.py) updates self.iterator.threshold and self.iterator.min_silence_samples immediately, allowing dynamic adjustment without pipeline restart.
What is the difference between min_speech_ms and min_speech_continuation_ms?
min_speech_ms (default 384 ms) defines the absolute minimum duration for a speech segment to be emitted as valid. min_speech_continuation_ms (default 192 ms) is a hysteresis value used specifically in realtime mode to determine whether a soft-ended turn should be reopened when new audio arrives. The continuation threshold must be less than or equal to the minimum speech duration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →