What Is the `--short_segment_merge_ms` VAD Parameter in Speech-to-Speech?

The --short_segment_merge_ms parameter defines an optional merge window that stitches together adjacent Voice Activity Detection (VAD) segments shorter than --min_speech_ms, preventing premature splitting of speech during brief pauses.

The huggingface/speech-to-speech repository provides a real-time speech-to-speech pipeline that relies on Voice Activity Detection (VAD) to segment audio streams. Understanding the --short_segment_merge_ms VAD parameter is essential when tuning aggressive VAD settings that might otherwise fragment continuous speech into disconnected chunks.

How --short_segment_merge_ms Works

When the VAD handler detects speech fragments shorter than the --min_speech_ms threshold, it faces a decision: discard the fragment as noise or retain it for potential merging. The --short_segment_merge_ms parameter controls this behavior by specifying a holding period in milliseconds.

If set to a value greater than 0, the handler retains short fragments for up to the specified duration. If another short fragment arrives within this merge window, the two fragments—including any silent gap between them—are concatenated into a single continuous segment. This prevents spurious turn splits caused by brief pauses in speech.

However, the implementation imposes a hard floor: fragments shorter than 100 ms are never held, regardless of the merge window setting. These ultra-short segments are immediately discarded as noise.

Configuration and Default Values

The parameter defaults to 0, which disables the merging behavior entirely. In this mode, any segment falling below --min_speech_ms is immediately discarded.

Enable this feature when using aggressive VAD configurations—such as very low --min_silence_ms values—where brief pauses in speech might trigger premature segment boundaries. The merge window provides robustness against micro-silences that occur naturally in conversational speech.

Implementation Details

The merging logic resides in the VADHandler class within src/speech_to_speech/VAD/vad_handler.py (lines 52–80). This handler manages the state machine for holding, timing, and concatenating eligible segments.

The CLI argument itself is defined in src/speech_to_speech/arguments_classes/vad_arguments.py (lines 76–80) as part of the VADHandlerArguments dataclass, which provides the interface between command-line flags and the handler configuration.

Usage Examples

You can enable short-segment merging via command-line flags or direct Python instantiation.

Using the CLI:

speech-to-speech \
  --mode realtime \
  --thresh 0.6 \
  --min_speech_ms 384 \
  --short_segment_merge_ms 300

Using the Python API:

from speech_to_speech.arguments_classes.vad_arguments import VADHandlerArguments
from speech_to_speech.s2s_pipeline import build_pipeline

vad_args = VADHandlerArguments(
    thresh=0.6,
    min_speech_ms=384,
    short_segment_merge_ms=300,
)

pipeline = build_pipeline(vad_handler_kwargs=vad_args)

Summary

  • --short_segment_merge_ms creates a holding window for VAD segments shorter than --min_speech_ms.
  • Default value is 0 (disabled); set to a positive integer (milliseconds) to enable merging.
  • 100 ms floor: Segments under 100 ms are always discarded as noise, even with merging enabled.
  • Gap tolerance: If the silence between two held fragments exceeds the merge window, the pending fragment is discarded.
  • Use case: Essential for aggressive VAD settings that use low --min_silence_ms values to prevent chopping continuous speech.

Frequently Asked Questions

What happens if I set --short_segment_merge_ms to 0?

When set to 0, the merge window is disabled. Any speech segment shorter than --min_speech_ms is immediately discarded rather than held for potential merging. This is the default behavior.

Why are segments shorter than 100 ms ignored even with merging enabled?

The implementation treats any detection under 100 ms as noise regardless of the merge window configuration. This hard limit prevents the system from attempting to merge transient artifacts or background clicks that the VAD might erroneously classify as speech.

How does this parameter differ from --min_silence_ms?

--min_silence_ms defines the minimum duration of silence required to split two speech segments, while --short_segment_merge_ms operates on the resulting segments themselves, allowing adjacent short segments to be recombined if they occur within the specified window. The former controls splitting; the latter controls post-split merging.

When should I enable short segment merging?

Enable this parameter when your VAD configuration produces fragmented output due to brief pauses in speech—such as when using a very low --min_silence_ms value to detect rapid speaker turns. It is particularly useful in real-time conversational settings where natural micropauses should not trigger separate processing segments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →