How to Swap STT Backends in the Hugging Face Speech-to-Speech Library
You swap STT backends by setting the stt field in ModuleArguments to your desired engine (e.g., whisper, faster-whisper, paraformer, or parakeet-tdt), which triggers the get_stt_handler factory in src/speech_to_speech/s2s_pipeline.py to instantiate the corresponding handler class.
The huggingface/speech-to-speech repository supports multiple Speech-to-Text (STT) engines through a modular handler architecture. To swap STT backends, you modify the stt parameter passed to the pipeline builder, either via command-line arguments or programmatic configuration objects.
Supported STT Backends and Handler Mapping
The pipeline supports six distinct STT engines, each mapped to a specific handler class:
stt Value |
Handler Class |
|---|---|
parakeet-tdt |
ParakeetTDTSTTHandler |
whisper |
WhisperSTTHandler |
whisper-mlx |
LightningWhisperSTTHandler |
mlx-audio-whisper |
MLXAudioWhisperSTTHandler |
faster-whisper |
FasterWhisperSTTHandler |
paraformer |
ParaformerSTTHandler |
Each backend requires its own argument dataclass (e.g., WhisperSTTHandlerArguments) defined in separate files under src/speech_to_speech/arguments_classes/.
How Backend Selection Works
The selection logic resides in the get_stt_handler function within src/speech_to_speech/s2s_pipeline.py (lines 69‑86 and 94‑108). When build_pipeline executes, it inspects the stt attribute of your ModuleArguments instance and instantiates the matching handler.
The ModuleArguments class is defined in src/speech_to_speech/arguments_classes/module_arguments.py and defaults to parakeet-tdt. Changing this field redirects the pipeline to load an alternative STT engine without modifying core pipeline code.
Swap Backends via Command Line
For CLI usage, pass the --stt flag followed by backend-specific configuration options. The argument parser automatically routes flags to the correct handler dataclass.
python -m speech_to_speech.demo.server \
--stt faster-whisper \
--faster_whisper_stt_model_name openai/whisper-large-v3 \
--faster_whisper_stt_device cuda \
--faster_whisper_stt_compute_type float16
Key flags explained:
--sttmaps toModuleArguments.sttand selects the handler type.- Backend‑specific flags (e.g.,
--whisper_model_name,--paraformer_device) populate the corresponding argument dataclass (e.g.,WhisperSTTHandlerArguments).
Swap Backends Programmatically
To switch backends in Python, import the appropriate argument classes and pass them to build_pipeline. Only the argument object matching the selected stt value is utilized; others serve as placeholders to satisfy the function signature.
from speech_to_speech.arguments_classes import (
ModuleArguments,
FasterWhisperSTTHandlerArguments,
WhisperSTTHandlerArguments,
ParaformerSTTHandlerArguments,
)
from speech_to_speech.s2s_pipeline import build_pipeline
# Configure the module to use Faster Whisper
module_args = ModuleArguments(stt="faster-whisper")
# Configure Faster Whisper specific settings
faster_args = FasterWhisperSTTHandlerArguments(
faster_whisper_stt_model_name="openai/whisper-large-v3",
faster_whisper_stt_device="cuda",
faster_whisper_stt_compute_type="float16",
)
# Build the pipeline - only faster_whisper_stt_handler_kwargs is used
pipeline = build_pipeline(
module_kwargs=module_args,
whisper_stt_handler_kwargs=WhisperSTTHandlerArguments(), # Placeholder
faster_whisper_stt_handler_kwargs=faster_args,
paraformer_stt_handler_kwargs=ParaformerSTTHandlerArguments(), # Placeholder
mlx_audio_whisper_stt_handler_kwargs=None,
parakeet_tdt_stt_handler_kwargs=None,
)
Backend Configuration File Reference
Each STT engine has a dedicated arguments file defining model names, devices, and generation parameters:
| Backend | Arguments File Path |
|---|---|
| Whisper | src/speech_to_speech/arguments_classes/whisper_stt_arguments.py |
| Faster Whisper | src/speech_to_speech/arguments_classes/faster_whisper_stt_arguments.py |
| Paraformer | src/speech_to_speech/arguments_classes/paraformer_stt_arguments.py |
| Parakeet TDT | src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py |
These dataclasses are unpacked via vars(<args>) when handlers initialize, allowing dynamic configuration injection.
Summary
- Backend selection occurs at startup via the
sttfield inModuleArguments. - The
get_stt_handlerfactory insrc/speech_to_speech/s2s_pipeline.pyinstantiates the appropriate handler class based on thesttvalue. - Six backends are supported:
parakeet-tdt,whisper,whisper-mlx,mlx-audio-whisper,faster-whisper, andparaformer. - Use
--stt <backend>for command-line switching or pass configured dataclasses tobuild_pipeline()for programmatic control. - Only the argument object matching the active
sttvalue is consumed; unused handler arguments are ignored but must be provided to satisfy thebuild_pipelinesignature.
Frequently Asked Questions
Can I switch STT backends without restarting the server?
No. The stt value in ModuleArguments is read only once during pipeline initialization in build_pipeline. To use a different backend, you must stop the current process and restart with a new ModuleArguments instance containing the desired stt flag.
What happens if I provide arguments for multiple STT backends?
The pipeline ignores argument objects that do not match the selected stt value. For example, if you set stt="faster-whisper" but pass configured objects for both WhisperSTTHandlerArguments and FasterWhisperSTTHandlerArguments, only the Faster Whisper configuration is utilized. However, you must still provide placeholder objects for unused handlers to satisfy the build_pipeline function signature.
Which STT backend is selected by default?
Parakeet TDT is the default STT engine. This is defined in src/speech_to_speech/arguments_classes/module_arguments.py where the stt field defaults to parakeet-tdt, loading the ParakeetTDTSTTHandler class automatically.
Do I need to install separate dependencies for each STT backend?
Yes. While the core pipeline can switch between handlers via configuration, each backend (Whisper, Faster Whisper, Paraformer, etc.) requires its own specific dependencies. Ensure you have installed the appropriate packages (e.g., faster-whisper, transformers, or paraformer) for your chosen backend before attempting to swap.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →