How the VAD Threshold (--thresh) Controls Turn-Taking in Speech-to-Speech
The --thresh parameter sets the probability threshold for the Voice Activity Detector (VAD), directly determining when a new user turn starts and how speculative turn reopening behaves during conversation.
The huggingface/speech-to-speech pipeline relies on Voice Activity Detection (VAD) to segment continuous audio into discrete user turns. The --thresh command-line argument controls the sensitivity of this detection, acting as the primary gatekeeper for turn-taking logic. Understanding how this threshold propagates from CLI arguments through to turn allocation helps optimize the system for different acoustic environments and conversational styles.
Threshold Propagation from CLI to VAD Iterator
Parsing the --thresh Argument
The threshold value originates in src/speech_to_speech/arguments_classes/vad_arguments.py (lines 5-11), where it is defined as a floating-point value with a default of 0.6. When the pipeline starts, this argument is parsed into VADHandlerArguments.thresh and forwarded to the handler initialization.
Handler Initialization
Inside src/speech_to_speech/VAD/vad_handler.py (lines 59-66 and 101-107), the VADHandler.setup method receives the threshold and instantiates the VADIterator:
self.iterator = VADIterator(
self.model,
threshold=thresh, # ← value from --thresh
sampling_rate=sample_rate,
min_silence_duration_ms=min_silence_ms,
speech_pad_ms=speech_pad_ms,
)
This VADIterator instance becomes the authoritative source for speech detection events throughout the session.
Trigger Logic and Turn Start Detection
Probability-Based Triggering
The core detection logic resides in src/speech_to_speech/VAD/vad_iterator.py (lines 131-138). For each audio chunk, the model computes a speech probability. The iterator triggers when this probability exceeds the configured threshold:
if (speech_prob >= self.threshold) and not self.triggered:
self.triggered = True
Lower thresholds (e.g., 0.3–0.4) cause the VAD to fire earlier, potentially detecting whispered speech or quiet starts but risking false triggers on background noise. Higher thresholds (e.g., 0.8+) require confident speech detection, preventing noise-induced turns but potentially missing brief or soft utterances.
Turn Allocation
When VADHandler.process detects that is_triggered_now has become True (derived from iterator.triggered), it evaluates whether the active speech duration satisfies the minimum required time before allocating a new turn. According to the source in src/speech_to_speech/VAD/vad_handler.py (lines 9-25), the handler computes effective_active_speech_duration_ms and compares it against active_speech_min_ms:
if is_triggered_now and not self._speech_started_emitted:
# … compute active_speech_duration_ms
active_speech_min_ms = self._active_speech_min_ms(effective_start_ms)
if effective_active_speech_duration_ms >= active_speech_min_ms:
turn_id, turn_revision, reopened = self._ensure_turn_for_speech_start(effective_start_ms)
Because the threshold determines when is_triggered_now becomes true, it implicitly controls when a turn starts. An aggressive (low) threshold may create turns on noise, while a conservative (high) threshold may delay turn start, potentially merging two short user utterances into a single turn.
Dynamic Threshold Updates at Runtime
In Realtime mode, the VAD threshold is not static. The client can send a RuntimeConfig containing a turn_detection.threshold field. The VADHandler._apply_runtime_turn_detection method (lines 169-172 in vad_handler.py) watches for this field and updates the iterator’s threshold on the fly:
if "threshold" in td:
self.iterator.threshold = td["threshold"]
This permits the server to adapt sensitivity during a session, such as lowering the threshold in quiet environments or raising it when background noise increases.
Interaction with Speculative Turn Reopening
Turn-taking behavior also involves speculative turns that remain reopenable for a short window defined by speculative_reopen_ms. Whether a turn can be reopened depends on _should_reopen_current_turn and the elapsed audio time.
As implemented in src/speech_to_speech/VAD/vad_handler.py (lines 20-22), the reopening logic checks:
if self._pending_reopen_candidate is not None or self._should_reopen_current_turn(start_ms):
# allow reopening ...
An overly aggressive threshold can cause premature turn starts that later get split or reopened, while an overly conservative threshold may suppress the opportunity to reopen a turn when the user continues speaking after a brief pause. The threshold effectively determines the temporal window available for speculative continuation (min_speech_continuation_ms applies only after a turn has started).
Practical Configuration Examples
Running with a Custom Threshold
Start the pipeline with a more sensitive VAD for quiet environments:
speech-to-speech \
--stt whisper \
--tts qwen3 \
--thresh 0.4 # more sensitive VAD
Programmatic Runtime Adjustment
Override the threshold dynamically during a session:
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.baseHandler import BaseHandler
from speech_to_speech.pipeline.handler_types import VADIn, VADOut
from threading import Event
# Create a handler with a custom threshold
handler = VADHandler()
handler.setup(
should_listen=Event(),
thresh=0.45, # custom threshold
sample_rate=16000,
min_silence_ms=64,
min_speech_ms=384,
)
# Later, during a realtime session, change the threshold
runtime_cfg = RuntimeConfig(
session=Session(
audio=AudioInput(turn_detection=TurnDetection(threshold=0.7))
)
)
handler.process((audio_chunk, runtime_cfg)) # threshold now 0.7
Debugging Turn Events
Inspect how threshold changes affect turn timing:
def print_events(vad_out: VADOut):
if isinstance(vad_out, SpeechStartedEvent):
print(f"Turn started → id={vad_out.turn_id} rev={vad_out.turn_revision}")
elif isinstance(vad_out, SpeechStoppedEvent):
print(f"Turn finished → id={vad_out.turn_id} rev={vad_out.turn_revision}")
# Hook the handler’s output queue
handler.text_output_queue = Queue()
while True:
for out in handler.process(audio_chunk):
print_events(out)
Summary
- The
--threshvalue flows fromvad_arguments.pytoVADHandler.setup, where it initializes theVADIteratorwith a specific probability threshold. - In
vad_iterator.py, the threshold determines whenself.triggeredbecomesTrue, which cascades into turn start detection via_ensure_turn_for_speech_startinvad_handler.py. - Lower thresholds increase sensitivity and risk false turn starts on noise, while higher thresholds delay detection and may merge short utterances.
- The threshold can be updated dynamically via
RuntimeConfigthrough_apply_runtime_turn_detection, allowing session-level adaptation. - Threshold settings interact with speculative turn reopening logic, affecting whether brief pauses split conversations or allow smooth continuation.
Frequently Asked Questions
What is the default VAD threshold in speech-to-speech?
The default value is 0.6, defined in src/speech_to_speech/arguments_classes/vad_arguments.py. This value provides a balance between detecting quiet speech and ignoring background noise for most environments.
How does lowering the VAD threshold affect barge-in detection?
Lowering the threshold makes the VAD trigger earlier, which can improve barge-in detection (interrupting the assistant) by catching the start of user speech sooner. However, if set too low (below 0.4), it may cause false barge-in events triggered by non-speech noise, requiring the speculative turn reopening logic to clean up spurious turns.
Can the VAD threshold be changed during an active conversation?
Yes. In Realtime mode, the client can send a RuntimeConfig payload containing a turn_detection.threshold field. The VADHandler._apply_runtime_turn_detection method (lines 169-172) detects this update and modifies self.iterator.threshold immediately, allowing dynamic adaptation without restarting the pipeline.
Why does a high VAD threshold cause missed turn starts?
A high threshold (e.g., 0.8 or above) requires the VAD model to produce high-confidence speech probabilities before triggering. Brief utterances, whispered speech, or the initial transient sounds of words may not reach this confidence level, causing the system to remain in a non-triggered state until the user speaks louder or longer, effectively delaying or missing the turn start event.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →