How to Swap STT Backends in the Hugging Face Speech-to-Speech Library

You swap STT backends by setting the stt field in ModuleArguments to your desired engine (e.g., whisper, faster-whisper, paraformer, or parakeet-tdt), which triggers the get_stt_handler factory in src/speech_to_speech/s2s_pipeline.py to instantiate the corresponding handler class.

The huggingface/speech-to-speech repository supports multiple Speech-to-Text (STT) engines through a modular handler architecture. To swap STT backends, you modify the stt parameter passed to the pipeline builder, either via command-line arguments or programmatic configuration objects.

Supported STT Backends and Handler Mapping

The pipeline supports six distinct STT engines, each mapped to a specific handler class:

stt Value Handler Class
parakeet-tdt ParakeetTDTSTTHandler
whisper WhisperSTTHandler
whisper-mlx LightningWhisperSTTHandler
mlx-audio-whisper MLXAudioWhisperSTTHandler
faster-whisper FasterWhisperSTTHandler
paraformer ParaformerSTTHandler

Each backend requires its own argument dataclass (e.g., WhisperSTTHandlerArguments) defined in separate files under src/speech_to_speech/arguments_classes/.

How Backend Selection Works

The selection logic resides in the get_stt_handler function within src/speech_to_speech/s2s_pipeline.py (lines 69‑86 and 94‑108). When build_pipeline executes, it inspects the stt attribute of your ModuleArguments instance and instantiates the matching handler.

The ModuleArguments class is defined in src/speech_to_speech/arguments_classes/module_arguments.py and defaults to parakeet-tdt. Changing this field redirects the pipeline to load an alternative STT engine without modifying core pipeline code.

Swap Backends via Command Line

For CLI usage, pass the --stt flag followed by backend-specific configuration options. The argument parser automatically routes flags to the correct handler dataclass.

python -m speech_to_speech.demo.server \
    --stt faster-whisper \
    --faster_whisper_stt_model_name openai/whisper-large-v3 \
    --faster_whisper_stt_device cuda \
    --faster_whisper_stt_compute_type float16

Key flags explained:

  • --stt maps to ModuleArguments.stt and selects the handler type.
  • Backend‑specific flags (e.g., --whisper_model_name, --paraformer_device) populate the corresponding argument dataclass (e.g., WhisperSTTHandlerArguments).

Swap Backends Programmatically

To switch backends in Python, import the appropriate argument classes and pass them to build_pipeline. Only the argument object matching the selected stt value is utilized; others serve as placeholders to satisfy the function signature.

from speech_to_speech.arguments_classes import (
    ModuleArguments,
    FasterWhisperSTTHandlerArguments,
    WhisperSTTHandlerArguments,
    ParaformerSTTHandlerArguments,
)
from speech_to_speech.s2s_pipeline import build_pipeline

# Configure the module to use Faster Whisper

module_args = ModuleArguments(stt="faster-whisper")

# Configure Faster Whisper specific settings

faster_args = FasterWhisperSTTHandlerArguments(
    faster_whisper_stt_model_name="openai/whisper-large-v3",
    faster_whisper_stt_device="cuda",
    faster_whisper_stt_compute_type="float16",
)

# Build the pipeline - only faster_whisper_stt_handler_kwargs is used

pipeline = build_pipeline(
    module_kwargs=module_args,
    whisper_stt_handler_kwargs=WhisperSTTHandlerArguments(),      # Placeholder

    faster_whisper_stt_handler_kwargs=faster_args,
    paraformer_stt_handler_kwargs=ParaformerSTTHandlerArguments(),  # Placeholder

    mlx_audio_whisper_stt_handler_kwargs=None,
    parakeet_tdt_stt_handler_kwargs=None,
)

Backend Configuration File Reference

Each STT engine has a dedicated arguments file defining model names, devices, and generation parameters:

Backend Arguments File Path
Whisper src/speech_to_speech/arguments_classes/whisper_stt_arguments.py
Faster Whisper src/speech_to_speech/arguments_classes/faster_whisper_stt_arguments.py
Paraformer src/speech_to_speech/arguments_classes/paraformer_stt_arguments.py
Parakeet TDT src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py

These dataclasses are unpacked via vars(<args>) when handlers initialize, allowing dynamic configuration injection.

Summary

  • Backend selection occurs at startup via the stt field in ModuleArguments.
  • The get_stt_handler factory in src/speech_to_speech/s2s_pipeline.py instantiates the appropriate handler class based on the stt value.
  • Six backends are supported: parakeet-tdt, whisper, whisper-mlx, mlx-audio-whisper, faster-whisper, and paraformer.
  • Use --stt <backend> for command-line switching or pass configured dataclasses to build_pipeline() for programmatic control.
  • Only the argument object matching the active stt value is consumed; unused handler arguments are ignored but must be provided to satisfy the build_pipeline signature.

Frequently Asked Questions

Can I switch STT backends without restarting the server?

No. The stt value in ModuleArguments is read only once during pipeline initialization in build_pipeline. To use a different backend, you must stop the current process and restart with a new ModuleArguments instance containing the desired stt flag.

What happens if I provide arguments for multiple STT backends?

The pipeline ignores argument objects that do not match the selected stt value. For example, if you set stt="faster-whisper" but pass configured objects for both WhisperSTTHandlerArguments and FasterWhisperSTTHandlerArguments, only the Faster Whisper configuration is utilized. However, you must still provide placeholder objects for unused handlers to satisfy the build_pipeline function signature.

Which STT backend is selected by default?

Parakeet TDT is the default STT engine. This is defined in src/speech_to_speech/arguments_classes/module_arguments.py where the stt field defaults to parakeet-tdt, loading the ParakeetTDTSTTHandler class automatically.

Do I need to install separate dependencies for each STT backend?

Yes. While the core pipeline can switch between handlers via configuration, each backend (Whisper, Faster Whisper, Paraformer, etc.) requires its own specific dependencies. Ensure you have installed the appropriate packages (e.g., faster-whisper, transformers, or paraformer) for your chosen backend before attempting to swap.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →