# How to Swap STT Backends in the Hugging Face Speech-to-Speech Library

> Easily swap STT backends like Whisper FasterWhisper and Parakeet in the Hugging Face Speech-to-Speech library Select your preferred engine with ModuleArguments for flexible ASR.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**You swap STT backends by setting the `stt` field in `ModuleArguments` to your desired engine (e.g., `whisper`, `faster-whisper`, `paraformer`, or `parakeet-tdt`), which triggers the `get_stt_handler` factory in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) to instantiate the corresponding handler class.**

The huggingface/speech-to-speech repository supports multiple Speech-to-Text (STT) engines through a modular handler architecture. To swap STT backends, you modify the `stt` parameter passed to the pipeline builder, either via command-line arguments or programmatic configuration objects.

## Supported STT Backends and Handler Mapping

The pipeline supports six distinct STT engines, each mapped to a specific handler class:

| `stt` Value | Handler Class |
|-------------|---------------|
| `parakeet-tdt` | `ParakeetTDTSTTHandler` |
| `whisper` | `WhisperSTTHandler` |
| `whisper-mlx` | `LightningWhisperSTTHandler` |
| `mlx-audio-whisper` | `MLXAudioWhisperSTTHandler` |
| `faster-whisper` | `FasterWhisperSTTHandler` |
| `paraformer` | `ParaformerSTTHandler` |

Each backend requires its own argument dataclass (e.g., `WhisperSTTHandlerArguments`) defined in separate files under `src/speech_to_speech/arguments_classes/`.

## How Backend Selection Works

The selection logic resides in the **`get_stt_handler`** function within **[`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)** (lines 69‑86 and 94‑108). When `build_pipeline` executes, it inspects the `stt` attribute of your `ModuleArguments` instance and instantiates the matching handler.

The **`ModuleArguments`** class is defined in **[`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)** and defaults to `parakeet-tdt`. Changing this field redirects the pipeline to load an alternative STT engine without modifying core pipeline code.

## Swap Backends via Command Line

For CLI usage, pass the `--stt` flag followed by backend-specific configuration options. The argument parser automatically routes flags to the correct handler dataclass.

```bash
python -m speech_to_speech.demo.server \
    --stt faster-whisper \
    --faster_whisper_stt_model_name openai/whisper-large-v3 \
    --faster_whisper_stt_device cuda \
    --faster_whisper_stt_compute_type float16

```

**Key flags explained:**
- `--stt` maps to `ModuleArguments.stt` and selects the handler type.
- Backend‑specific flags (e.g., `--whisper_model_name`, `--paraformer_device`) populate the corresponding argument dataclass (e.g., `WhisperSTTHandlerArguments`).

## Swap Backends Programmatically

To switch backends in Python, import the appropriate argument classes and pass them to `build_pipeline`. Only the argument object matching the selected `stt` value is utilized; others serve as placeholders to satisfy the function signature.

```python
from speech_to_speech.arguments_classes import (
    ModuleArguments,
    FasterWhisperSTTHandlerArguments,
    WhisperSTTHandlerArguments,
    ParaformerSTTHandlerArguments,
)
from speech_to_speech.s2s_pipeline import build_pipeline

# Configure the module to use Faster Whisper

module_args = ModuleArguments(stt="faster-whisper")

# Configure Faster Whisper specific settings

faster_args = FasterWhisperSTTHandlerArguments(
    faster_whisper_stt_model_name="openai/whisper-large-v3",
    faster_whisper_stt_device="cuda",
    faster_whisper_stt_compute_type="float16",
)

# Build the pipeline - only faster_whisper_stt_handler_kwargs is used

pipeline = build_pipeline(
    module_kwargs=module_args,
    whisper_stt_handler_kwargs=WhisperSTTHandlerArguments(),      # Placeholder

    faster_whisper_stt_handler_kwargs=faster_args,
    paraformer_stt_handler_kwargs=ParaformerSTTHandlerArguments(),  # Placeholder

    mlx_audio_whisper_stt_handler_kwargs=None,
    parakeet_tdt_stt_handler_kwargs=None,
)

```

## Backend Configuration File Reference

Each STT engine has a dedicated arguments file defining model names, devices, and generation parameters:

| Backend | Arguments File Path |
|---------|-------------------|
| **Whisper** | [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py) |
| **Faster Whisper** | [`src/speech_to_speech/arguments_classes/faster_whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/faster_whisper_stt_arguments.py) |
| **Paraformer** | [`src/speech_to_speech/arguments_classes/paraformer_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/paraformer_stt_arguments.py) |
| **Parakeet TDT** | [`src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/parakeet_tdt_arguments.py) |

These dataclasses are unpacked via `vars(<args>)` when handlers initialize, allowing dynamic configuration injection.

## Summary

- **Backend selection** occurs at startup via the `stt` field in `ModuleArguments`.
- The **`get_stt_handler`** factory in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) instantiates the appropriate handler class based on the `stt` value.
- **Six backends** are supported: `parakeet-tdt`, `whisper`, `whisper-mlx`, `mlx-audio-whisper`, `faster-whisper`, and `paraformer`.
- Use `--stt <backend>` for command-line switching or pass configured dataclasses to `build_pipeline()` for programmatic control.
- Only the argument object matching the active `stt` value is consumed; unused handler arguments are ignored but must be provided to satisfy the `build_pipeline` signature.

## Frequently Asked Questions

### Can I switch STT backends without restarting the server?

No. The `stt` value in `ModuleArguments` is read only once during pipeline initialization in `build_pipeline`. To use a different backend, you must stop the current process and restart with a new `ModuleArguments` instance containing the desired `stt` flag.

### What happens if I provide arguments for multiple STT backends?

The pipeline ignores argument objects that do not match the selected `stt` value. For example, if you set `stt="faster-whisper"` but pass configured objects for both `WhisperSTTHandlerArguments` and `FasterWhisperSTTHandlerArguments`, only the Faster Whisper configuration is utilized. However, you must still provide placeholder objects for unused handlers to satisfy the `build_pipeline` function signature.

### Which STT backend is selected by default?

**Parakeet TDT** is the default STT engine. This is defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) where the `stt` field defaults to `parakeet-tdt`, loading the `ParakeetTDTSTTHandler` class automatically.

### Do I need to install separate dependencies for each STT backend?

Yes. While the core pipeline can switch between handlers via configuration, each backend (Whisper, Faster Whisper, Paraformer, etc.) requires its own specific dependencies. Ensure you have installed the appropriate packages (e.g., `faster-whisper`, `transformers`, or `paraformer`) for your chosen backend before attempting to swap.