# What Is the `--short_segment_merge_ms` VAD Parameter in Speech-to-Speech?

> Discover the `--short_segment_merge_ms` VAD parameter in speech-to-speech. Learn how it merges brief VAD segments to prevent speech splitting during short pauses.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-11

---

**The `--short_segment_merge_ms` parameter defines an optional merge window that stitches together adjacent Voice Activity Detection (VAD) segments shorter than `--min_speech_ms`, preventing premature splitting of speech during brief pauses.**

The `huggingface/speech-to-speech` repository provides a real-time speech-to-speech pipeline that relies on Voice Activity Detection (VAD) to segment audio streams. Understanding the `--short_segment_merge_ms` VAD parameter is essential when tuning aggressive VAD settings that might otherwise fragment continuous speech into disconnected chunks.

## How `--short_segment_merge_ms` Works

When the VAD handler detects speech fragments shorter than the `--min_speech_ms` threshold, it faces a decision: discard the fragment as noise or retain it for potential merging. The `--short_segment_merge_ms` parameter controls this behavior by specifying a holding period in milliseconds.

If set to a value greater than `0`, the handler retains short fragments for up to the specified duration. If another short fragment arrives within this merge window, the two fragments—including any silent gap between them—are concatenated into a single continuous segment. This prevents spurious turn splits caused by brief pauses in speech.

However, the implementation imposes a hard floor: **fragments shorter than 100 ms** are never held, regardless of the merge window setting. These ultra-short segments are immediately discarded as noise.

## Configuration and Default Values

The parameter defaults to `0`, which disables the merging behavior entirely. In this mode, any segment falling below `--min_speech_ms` is immediately discarded.

Enable this feature when using aggressive VAD configurations—such as very low `--min_silence_ms` values—where brief pauses in speech might trigger premature segment boundaries. The merge window provides robustness against micro-silences that occur naturally in conversational speech.

## Implementation Details

The merging logic resides in the `VADHandler` class within [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) (lines 52–80). This handler manages the state machine for holding, timing, and concatenating eligible segments.

The CLI argument itself is defined in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py) (lines 76–80) as part of the `VADHandlerArguments` dataclass, which provides the interface between command-line flags and the handler configuration.

## Usage Examples

You can enable short-segment merging via command-line flags or direct Python instantiation.

Using the CLI:

```bash
speech-to-speech \
  --mode realtime \
  --thresh 0.6 \
  --min_speech_ms 384 \
  --short_segment_merge_ms 300

```

Using the Python API:

```python
from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.s2s_pipeline import build_pipeline

vad_args = VADHandlerArguments(
    thresh=0.6,
    min_speech_ms=384,
    short_segment_merge_ms=300,
)

pipeline = build_pipeline(vad_handler_kwargs=vad_args)

```

## Summary

- **`--short_segment_merge_ms`** creates a holding window for VAD segments shorter than `--min_speech_ms`.
- **Default value** is `0` (disabled); set to a positive integer (milliseconds) to enable merging.
- **100 ms floor**: Segments under 100 ms are always discarded as noise, even with merging enabled.
- **Gap tolerance**: If the silence between two held fragments exceeds the merge window, the pending fragment is discarded.
- **Use case**: Essential for aggressive VAD settings that use low `--min_silence_ms` values to prevent chopping continuous speech.

## Frequently Asked Questions

### What happens if I set `--short_segment_merge_ms` to 0?

When set to `0`, the merge window is disabled. Any speech segment shorter than `--min_speech_ms` is immediately discarded rather than held for potential merging. This is the default behavior.

### Why are segments shorter than 100 ms ignored even with merging enabled?

The implementation treats any detection under 100 ms as noise regardless of the merge window configuration. This hard limit prevents the system from attempting to merge transient artifacts or background clicks that the VAD might erroneously classify as speech.

### How does this parameter differ from `--min_silence_ms`?

`--min_silence_ms` defines the minimum duration of silence required to split two speech segments, while `--short_segment_merge_ms` operates on the resulting segments themselves, allowing adjacent short segments to be recombined if they occur within the specified window. The former controls splitting; the latter controls post-split merging.

### When should I enable short segment merging?

Enable this parameter when your VAD configuration produces fragmented output due to brief pauses in speech—such as when using a very low `--min_silence_ms` value to detect rapid speaker turns. It is particularly useful in real-time conversational settings where natural micropauses should not trigger separate processing segments.