# How the VAD Threshold (--thresh) Controls Turn-Taking in Speech-to-Speech

> Learn how the VAD threshold --thresh controls turn-taking in speech-to-speech conversations. Adjust this setting to manage when user turns begin and how turn reopening behaves.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-11

---

**The `--thresh` parameter sets the probability threshold for the Voice Activity Detector (VAD), directly determining when a new user turn starts and how speculative turn reopening behaves during conversation.**

The huggingface/speech-to-speech pipeline relies on Voice Activity Detection (VAD) to segment continuous audio into discrete user turns. The `--thresh` command-line argument controls the sensitivity of this detection, acting as the primary gatekeeper for turn-taking logic. Understanding how this threshold propagates from CLI arguments through to turn allocation helps optimize the system for different acoustic environments and conversational styles.

## Threshold Propagation from CLI to VAD Iterator

### Parsing the --thresh Argument

The threshold value originates in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py) (lines 5-11), where it is defined as a floating-point value with a default of `0.6`. When the pipeline starts, this argument is parsed into `VADHandlerArguments.thresh` and forwarded to the handler initialization.

### Handler Initialization

Inside [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) (lines 59-66 and 101-107), the `VADHandler.setup` method receives the threshold and instantiates the `VADIterator`:

```python
self.iterator = VADIterator(
    self.model,
    threshold=thresh,                     # ← value from --thresh

    sampling_rate=sample_rate,
    min_silence_duration_ms=min_silence_ms,
    speech_pad_ms=speech_pad_ms,
)

```

This `VADIterator` instance becomes the authoritative source for speech detection events throughout the session.

## Trigger Logic and Turn Start Detection

### Probability-Based Triggering

The core detection logic resides in [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py) (lines 131-138). For each audio chunk, the model computes a speech probability. The iterator triggers when this probability exceeds the configured threshold:

```python
if (speech_prob >= self.threshold) and not self.triggered:
    self.triggered = True

```

**Lower thresholds** (e.g., 0.3–0.4) cause the VAD to fire earlier, potentially detecting whispered speech or quiet starts but risking false triggers on background noise. **Higher thresholds** (e.g., 0.8+) require confident speech detection, preventing noise-induced turns but potentially missing brief or soft utterances.

### Turn Allocation

When `VADHandler.process` detects that `is_triggered_now` has become `True` (derived from `iterator.triggered`), it evaluates whether the **active speech duration** satisfies the minimum required time before allocating a new turn. According to the source in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) (lines 9-25), the handler computes `effective_active_speech_duration_ms` and compares it against `active_speech_min_ms`:

```python
if is_triggered_now and not self._speech_started_emitted:
    # … compute active_speech_duration_ms

    active_speech_min_ms = self._active_speech_min_ms(effective_start_ms)
    if effective_active_speech_duration_ms >= active_speech_min_ms:
        turn_id, turn_revision, reopened = self._ensure_turn_for_speech_start(effective_start_ms)

```

Because the threshold determines when `is_triggered_now` becomes true, it implicitly controls **when a turn starts**. An aggressive (low) threshold may create turns on noise, while a conservative (high) threshold may delay turn start, potentially merging two short user utterances into a single turn.

## Dynamic Threshold Updates at Runtime

In Realtime mode, the VAD threshold is not static. The client can send a `RuntimeConfig` containing a `turn_detection.threshold` field. The `VADHandler._apply_runtime_turn_detection` method (lines 169-172 in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py)) watches for this field and updates the iterator’s threshold on the fly:

```python
if "threshold" in td:
    self.iterator.threshold = td["threshold"]

```

This permits the server to adapt sensitivity during a session, such as lowering the threshold in quiet environments or raising it when background noise increases.

## Interaction with Speculative Turn Reopening

Turn-taking behavior also involves **speculative turns** that remain reopenable for a short window defined by `speculative_reopen_ms`. Whether a turn can be reopened depends on `_should_reopen_current_turn` and the elapsed audio time. 

As implemented in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) (lines 20-22), the reopening logic checks:

```python
if self._pending_reopen_candidate is not None or self._should_reopen_current_turn(start_ms):
    # allow reopening ...

```

An overly aggressive threshold can cause *premature* turn starts that later get split or reopened, while an overly conservative threshold may suppress the opportunity to reopen a turn when the user continues speaking after a brief pause. The threshold effectively determines the temporal window available for speculative continuation (`min_speech_continuation_ms` applies only after a turn has started).

## Practical Configuration Examples

### Running with a Custom Threshold

Start the pipeline with a more sensitive VAD for quiet environments:

```bash
speech-to-speech \
  --stt whisper \
  --tts qwen3 \
  --thresh 0.4   # more sensitive VAD

```

### Programmatic Runtime Adjustment

Override the threshold dynamically during a session:

```python
from speech_to_speech.VAD.vad_handler import VADHandler
from speech_to_speech.baseHandler import BaseHandler
from speech_to_speech.pipeline.handler_types import VADIn, VADOut
from threading import Event

# Create a handler with a custom threshold

handler = VADHandler()
handler.setup(
    should_listen=Event(),
    thresh=0.45,               # custom threshold

    sample_rate=16000,
    min_silence_ms=64,
    min_speech_ms=384,
)

# Later, during a realtime session, change the threshold

runtime_cfg = RuntimeConfig(
    session=Session(
        audio=AudioInput(turn_detection=TurnDetection(threshold=0.7))
    )
)
handler.process((audio_chunk, runtime_cfg))  # threshold now 0.7

```

### Debugging Turn Events

Inspect how threshold changes affect turn timing:

```python
def print_events(vad_out: VADOut):
    if isinstance(vad_out, SpeechStartedEvent):
        print(f"Turn started → id={vad_out.turn_id} rev={vad_out.turn_revision}")
    elif isinstance(vad_out, SpeechStoppedEvent):
        print(f"Turn finished → id={vad_out.turn_id} rev={vad_out.turn_revision}")

# Hook the handler’s output queue

handler.text_output_queue = Queue()
while True:
    for out in handler.process(audio_chunk):
        print_events(out)

```

## Summary

- The `--thresh` value flows from [`vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_arguments.py) to `VADHandler.setup`, where it initializes the `VADIterator` with a specific probability threshold.
- In [`vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_iterator.py), the threshold determines when `self.triggered` becomes `True`, which cascades into turn start detection via `_ensure_turn_for_speech_start` in [`vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_handler.py).
- Lower thresholds increase sensitivity and risk false turn starts on noise, while higher thresholds delay detection and may merge short utterances.
- The threshold can be updated dynamically via `RuntimeConfig` through `_apply_runtime_turn_detection`, allowing session-level adaptation.
- Threshold settings interact with speculative turn reopening logic, affecting whether brief pauses split conversations or allow smooth continuation.

## Frequently Asked Questions

### What is the default VAD threshold in speech-to-speech?

The default value is **0.6**, defined in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py). This value provides a balance between detecting quiet speech and ignoring background noise for most environments.

### How does lowering the VAD threshold affect barge-in detection?

Lowering the threshold makes the VAD trigger earlier, which can improve barge-in detection (interrupting the assistant) by catching the start of user speech sooner. However, if set too low (below 0.4), it may cause false barge-in events triggered by non-speech noise, requiring the speculative turn reopening logic to clean up spurious turns.

### Can the VAD threshold be changed during an active conversation?

Yes. In Realtime mode, the client can send a `RuntimeConfig` payload containing a `turn_detection.threshold` field. The `VADHandler._apply_runtime_turn_detection` method (lines 169-172) detects this update and modifies `self.iterator.threshold` immediately, allowing dynamic adaptation without restarting the pipeline.

### Why does a high VAD threshold cause missed turn starts?

A high threshold (e.g., 0.8 or above) requires the VAD model to produce high-confidence speech probabilities before triggering. Brief utterances, whispered speech, or the initial transient sounds of words may not reach this confidence level, causing the system to remain in a non-triggered state until the user speaks louder or longer, effectively delaying or missing the turn start event.