# Are There Pre-Trained Speech-to-Speech Models Available?

> Discover readily available pre-trained speech-to-speech models from huggingface. Access and utilize multiple models directly with automatic downloads from the Hugging Face Hub.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-08-01

---

**Yes, the huggingface/speech-to-speech library ships with ready-to-use wrappers for multiple pre-trained speech-to-speech models that download automatically from the Hugging Face Hub.**

The huggingface/speech-to-speech repository provides a unified **S2SPipeline** that orchestrates end-to-end voice processing using pre-trained speech-to-speech models. You can initialize production-ready pipelines for text-to-speech, speech-to-text, and full speech-to-speech conversion without manual model downloads or complex configuration.

## How the Pipeline Loads Pre-Trained Models

The architecture centers on the `S2SPipeline` class in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py), which acts as a high-level orchestrator. When you call `from_pretrained()`, the pipeline inspects the provided arguments, resolves platform-specific variants (such as Apple Silicon MLX quantization), and delegates to specialized handler classes that invoke `from_pretrained()` on the underlying model architectures.

**Argument dataclasses** expose model identifiers through CLI-friendly interfaces. For example, `Qwen3TTSArguments` in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) defines the `qwen3_tts_model_name` parameter with a default value of `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`.

**Handler classes** implement the actual loading logic. The `Qwen3TTSHandler` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) supports both Faster-Transformer and MLX backends, automatically appending `-6bit` suffixes for MLX quantized variants when needed.

## Supported Pre-Trained Model Families

The library supports diverse pre-trained speech models across the STT → LLM → TTS stack.

### Qwen3-TTS Models

The primary TTS backend uses **Qwen3-TTS** models, loaded via `Qwen3TTSHandler`. The handler resolves model identifiers and initializes either `FasterQwen3TTS` for standard CUDA/CPU inference or MLX-optimized variants for Apple Silicon. Default models include `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` and MLX community variants like `mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16`.

### Whisper STT Models

For speech-to-text, the pipeline integrates **Whisper** models through `WhisperSTTArguments` (defined in [`src/speech_to_speech/arguments_classes/whisper_stt_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/whisper_stt_arguments.py)) and `WhisperSTTHandler` (implemented in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py)). The handler loads models using `AutoProcessor.from_pretrained()` and `AutoModelForSpeechSeq2Seq.from_pretrained()`, with defaults pointing to `distil-whisper/distil-large-v3`.

### Alternative TTS Backends

Beyond Qwen3-TTS, the library exposes handlers for **Parler-TTS**, **Parakeet**, **Paraformer**, **Facebook-MMS**, **Kokoro**, **Pocket TTS**, and **ChatTTS**. Each backend provides its own arguments class (e.g., `ParlerTTSArguments` in [`src/speech_to_speech/arguments_classes/parler_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/parler_tts_arguments.py)) and handler implementation, following the same `from_pretrained` pattern.

## Practical Code Examples

Initialize a pipeline with default pre-trained models:

```python
from speech_to_speech.s2s_pipeline import S2SPipeline

pipeline = S2SPipeline.from_pretrained()
audio_output = pipeline.run_text("Hello, how are you?")

```

Specify a specific MLX-optimized model for Apple Silicon:

```python
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments

args = Qwen3TTSArguments(
    qwen3_tts_model_name="mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16"
)
pipeline = S2SPipeline.from_pretrained(qwen3_tts_args=args)
audio_output = pipeline.run_text("Bonjour, je suis un modèle TTS.")

```

Use the Parler-TTS backend instead:

```python
from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.parler_tts_arguments import ParlerTTSArguments

args = ParlerTTSArguments(parler_tts_model_name="parler-tts/parler-mini-v1-jenny")
pipeline = S2SPipeline.from_pretrained(parler_tts_args=args)
audio_output = pipeline.run_text("Testing the Parler TTS model.")

```

All examples download model weights automatically on first execution.

## Summary

- The **S2SPipeline** in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) provides unified access to pre-trained speech-to-speech models.
- **Qwen3TTSHandler** and **WhisperSTTHandler** manage model instantiation for TTS and STT components respectively.
- Models download automatically from the Hugging Face Hub when calling `from_pretrained()`.
- Apple Silicon users receive optimized MLX variants through automatic suffix resolution (e.g., `-6bit`).

## Frequently Asked Questions

### What pre-trained models are included by default?

The pipeline defaults to `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` for text-to-speech and `distil-whisper/distil-large-v3` for speech-to-text. These initialize automatically when you call `S2SPipeline.from_pretrained()` without arguments.

### Can I use custom Hugging Face models?

Yes. Pass any valid Hugging Face Hub model identifier to the appropriate arguments class, such as `Qwen3TTSArguments(qwen3_tts_model_name="your-username/your-model")`. The handler validates the identifier and downloads weights via the standard `from_pretrained()` mechanism.

### Does the library support quantized models for Apple Silicon?

Yes. When running on Apple Silicon, `Qwen3TTSHandler` detects the platform and automatically resolves MLX-quantized variants by appending `-6bit` to the model name, or you can explicitly specify MLX community models like `mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16`.

### How do I switch between different TTS backends?

Import the specific arguments class for your desired backend (e.g., `ParlerTTSArguments` or `ParakeetTDTArguments`) and pass the instance to `S2SPipeline.from_pretrained()` using the appropriate parameter name (e.g., `parler_tts_args` or `parakeet_tdt_args`).