How to Configure VAD Parameters in Hugging Face Speech-to-Speech: Thresholds, Silence Detection, and Realtime Settings

Configure VAD parameters by instantiating VADHandlerArguments with your desired thresholds and timings, then pass it to SpeechToSpeechPipeline, or update settings dynamically via the OpenAI-compatible realtime API's session.update event.

The Hugging Face speech-to-speech repository wraps the Silero VAD model inside VADHandler to detect voice activity in audio streams. All user-controllable VAD options are centralized in the VADHandlerArguments dataclass and applied through the handler's setup() method, giving you precise control over speech boundaries, silence detection, and realtime processing behavior.

Understanding VADHandlerArguments

The VADHandlerArguments dataclass in src/speech_to_speech/arguments_classes/vad_arguments.py exposes every tunable knob for voice activity detection. When you instantiate SpeechToSpeechPipeline, these arguments are forwarded to VADHandler.setup() in src/speech_to_speech/VAD/vad_handler.py, which initializes the internal VADIterator.

Speech Detection Thresholds

The thresh parameter (defined at vad_arguments.py#L6-L10) controls the confidence level required for the Silero model to consider audio as speech. It accepts float values between 0.0 and 1.0, with a default of 0.6. Higher values reduce false positives but may miss quiet speech, while lower values increase sensitivity to background noise.

Speech and Silence Duration Controls

Several parameters govern the minimum and maximum durations for speech segments and silence gaps:

  • min_silence_ms (default 64 ms): Minimum continuous silence that triggers a segment split. Defined at vad_arguments.py#L18-L22.
  • min_speech_ms (default 384 ms): Minimum length for a valid speech utterance. Shorter fragments are held for merging or discarded. Defined at vad_arguments.py#L24-L28.
  • min_speech_continuation_ms (default 192 ms): Hysteresis threshold for reopening a speech turn in realtime mode. Must be less than or equal to min_speech_ms. Defined at vad_arguments.py#L30-L34.
  • max_speech_ms (default ∞): Hard limit on segment duration; forces a split when exceeded. Defined at vad_arguments.py#L36-L40.
  • speech_pad_ms (default 500 ms in arguments, overridable to 30 ms in setup): Amount of audio prepended to each segment to capture speech onsets. Defined at vad_arguments.py#L42-L46.

Realtime Processing and Enhancement Options

For streaming applications, the following settings control progressive transcription and audio preprocessing:

  • enable_realtime_transcription (default False): Emits audio chunks while the user is still speaking. Defined at vad_arguments.py#L54-L57.
  • realtime_processing_pause (default 0.5 s): Base interval between progressive chunks, automatically scaled by speech length. Defined at vad_arguments.py#L58-L62.
  • speculative_reopen_ms (default 1000 ms) and unanswered_reopen_ms (default 7000 ms): Duration that a soft-ended turn remains reopenable. Defined at vad_arguments.py#L64-L74.
  • short_segment_merge_ms (default 0): Window allowing tiny fragments to be stitched together. Useful when min_silence_ms is set very low. Defined at vad_arguments.py#L76-L80.
  • audio_enhancement (default False): Enables DeepFilterNet noise reduction. Requires the optional df package. Defined at vad_arguments.py#L48-L53.

Initializing VAD Parameters at Startup

Pass a configured VADHandlerArguments instance when constructing your pipeline. The setup() method receives these values and instantiates the VADIterator with the specified threshold, sampling rate, and silence duration:

from speech_to_speech.pipeline import SpeechToSpeechPipeline
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments

vad_args = VADHandlerArguments(
    thresh=0.75,               # stricter detection

    min_silence_ms=100,        # longer silence before split

    min_speech_ms=250,         # accept shorter utterances

    speech_pad_ms=200,         # keep only 200 ms of pre-speech context

    audio_enhancement=True,    # enable noise reduction (requires df)

)

pipeline = SpeechToSpeechPipeline(
    vad_handler_args=vad_args,
    # … other arguments like model, tokenizer, etc.

)

Inside VADHandler.setup() (lines 59–77 of vad_handler.py), these arguments populate the VADIterator:

self.iterator = VADIterator(
    self.model,
    threshold=thresh,
    sampling_rate=sample_rate,
    min_silence_duration_ms=min_silence_ms,
    speech_pad_ms=speech_pad_ms,
)

Updating VAD Configuration at Runtime

When using the OpenAI-compatible realtime endpoint, clients can adjust VAD parameters without restarting the server. The handler watches for session.update events and applies changes via _apply_runtime_turn_detection (lines 45–74 of vad_handler.py).

Send a JSON message over the websocket to modify the threshold or silence duration:


# Assuming you have a websocket session `session`

session.send({
    "type": "session.update",
    "session": {
        "audio": {
            "input": {
                "turn_detection": {
                    "type": "server_vad",
                    "threshold": 0.8,               # raise confidence

                    "silence_duration_ms": 120       # lengthen silence window

                }
            }
        }
    }
})

The handler updates self.iterator.threshold and self.iterator.min_silence_samples immediately upon receiving the event.

Handling Edge Cases: Short Segments and Audio Enhancement

Merging Short Speech Segments

When min_silence_ms is set aggressively low, you may generate many tiny fragments. Enable stitching by setting short_segment_merge_ms to a positive value:

vad_args = VADHandlerArguments(
    short_segment_merge_ms=150,   # allow fragments within 150 ms to be merged

)

pipeline = SpeechToSpeechPipeline(vad_handler_args=vad_args)

With this setting, the handler invokes _hold_short_segment (lines 107–115) and _merge_pending_short_segment (lines 83–106) to combine fragments before discarding them.

Enabling Progressive Realtime Transcription

For low-latency streaming, enable progressive audio release:

vad_args = VADHandlerArguments(
    enable_realtime_transcription=True,
    realtime_processing_pause=0.3,   # emit chunks every ~300 ms (scaled automatically)

)

pipeline = SpeechToSpeechPipeline(vad_handler_args=vad_args)

When enabled, the handler enters the _process_realtime branch (beginning at line 69 of vad_handler.py), yielding VADAudio objects in "progressive" mode while the user is still speaking.

Summary

  • Configuration Source: All VAD defaults live in VADHandlerArguments (src/speech_to_speech/arguments_classes/vad_arguments.py), while runtime logic resides in VADHandler (src/speech_to_speech/VAD/vad_handler.py).
  • Threshold Control: Adjust thresh (0.0–1.0) to balance between false positives and missed speech.
  • Segment Boundaries: Use min_silence_ms, min_speech_ms, and speech_pad_ms to define how audio is split and padded.
  • Realtime Updates: Modify threshold and silence_duration_ms on-the-fly via the OpenAI-compatible API's session.update event, processed by _apply_runtime_turn_detection.
  • Advanced Features: Enable audio_enhancement for noise reduction, short_segment_merge_ms for fragment stitching, and enable_realtime_transcription for progressive streaming.

Frequently Asked Questions

What is the default VAD threshold in the speech-to-speech pipeline?

The default thresh value is 0.6, defined in src/speech_to_speech/arguments_classes/vad_arguments.py. This means the Silero model must output a probability of at least 60% for the audio to be classified as speech. You can override this at initialization or update it at runtime via the realtime API.

How do I enable realtime transcription with custom VAD timing?

Set enable_realtime_transcription=True in your VADHandlerArguments and adjust realtime_processing_pause to control the base interval between chunks. For example, setting realtime_processing_pause=0.3 yields emissions approximately every 300 milliseconds, scaled automatically by the handler based on speech length.

Can I change VAD settings without restarting the server?

Yes. When using the OpenAI-compatible realtime endpoint, send a session.update event with the turn_detection payload. The VADHandler._apply_runtime_turn_detection method (lines 45–74 of vad_handler.py) updates self.iterator.threshold and self.iterator.min_silence_samples immediately, allowing dynamic adjustment without pipeline restart.

What is the difference between min_speech_ms and min_speech_continuation_ms?

min_speech_ms (default 384 ms) defines the absolute minimum duration for a speech segment to be emitted as valid. min_speech_continuation_ms (default 192 ms) is a hysteresis value used specifically in realtime mode to determine whether a soft-ended turn should be reopened when new audio arrives. The continuation threshold must be less than or equal to the minimum speech duration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →