# Where to Find Pre-Trained Speech-to-Speech Models: Hugging Face Hub Integration Guide

> Discover pre-trained speech-to-speech models on Hugging Face Hub. Easily integrate VAD, STT, LLM, and TTS components using the CLI or Python API. Download weights automatically.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: getting-started
- Published: 2026-07-07

---

**Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and can be referenced directly by repository ID in the CLI or Python API to automatically download and cache weights for VAD, STT, LLM, and TTS components.**

The `huggingface/speech-to-speech` repository provides a production-ready pipeline that ships with pre-trained speech-to-speech models for every stage of voice processing. According to the source code, all default model weights are hosted on the Hugging Face Hub and resolve automatically when you reference them by name, triggering download and local caching on first use.

## Default Pre-Trained Models Available Out-of-the-Box

The pipeline includes ready-to-use models for each component of the speech-to-speech stack.

### Voice Activity Detection (VAD)

The pipeline uses **Silero VAD v5** (`snakers4/silero-vad`) for voice activity detection. This model is loaded automatically by the VAD handler in [`src/speech_to_speech/VAD/silero_vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/silero_vad_handler.py) without requiring manual configuration.

### Speech-to-Text (STT)

For speech recognition, the default is **Parakeet TDT 0.6B v3**. The handler at [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) implements platform-specific resolution: it selects `mlx-community/parakeet-tdt-0.6b-v3` on macOS and `nvidia/parakeet-tdt-0.6b-v3` on other platforms (lines 144-152).

### Large Language Model (LLM)

The default language model is **OpenAI GPT-5.4-mini** (`openai/gpt-5.4-mini`), managed by [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py). This component handles the text generation stage between speech recognition and synthesis.

### Text-to-Speech (TTS)

The default TTS backend uses **Qwen3-TTS 12Hz 1.7B CustomVoice** (`Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`). The handler in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) includes automatic quantization logic in the `_resolve_mlx_model_name` function (lines 286-293), which appends a `-6bit` suffix to the model name if not already specified.

## Optional TTS Backends and Alternative Pre-Trained Models

The pipeline supports several alternative TTS models that can be swapped via CLI flags:

- **Kokoro-82M** (`hexgrad/Kokoro-82M`) – Implemented in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py)
- **Pocket TTS** (`kyutai-labs/pocket-tts`) – Implemented in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) with voice cloning capabilities
- **ChatTTS** (`2noise/ChatTTS`) – Implemented in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) for English and Chinese synthesis
- **MMS TTS** (`facebook/mms-tts`) – Implemented in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) for multilingual support

## How Model Resolution Works in the Source Code

The repository implements automatic model resolution logic that determines whether to use MLX, GGML, or standard Transformers backends based on the current platform.

In [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the `_resolve_mlx_model_name` function handles quantization defaults by adding a `-6bit` quantization suffix when the model name does not already specify one (lines 286-293). Similarly, the STT handler in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) selects between MLX and NVIDIA variants based on the operating system (lines 144-152).

When you reference a model by its Hugging Face Hub identifier, the library downloads the weights to the local cache on first use and loads them into the appropriate handler.

## CLI Examples for Using Pre-Trained Models

You can override default models using the `--*_model_name` flags or switch backends with `--tts`, `--stt`, and `--llm` flags.

### Run the Pipeline with Default Models

Start the WebSocket server using the default pre-trained models (Parakeet TDT + Qwen3-TTS + OpenAI LLM):

```bash
export OPENAI_API_KEY=YOUR_KEY
speech-to-speech

```

### Use Kokoro TTS Instead of Qwen3

Switch to the Kokoro-82M model by selecting the Kokoro handler and specifying the repository ID:

```bash
speech-to-speech \
  --tts kokoro \
  --kokoro_model_name hexgrad/Kokoro-82M \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

```

### Configure Pocket TTS with a Specific Voice

Use the Pocket TTS backend with a voice preset and CPU inference:

```bash
speech-to-speech \
  --tts pocket \
  --pocket_tts_voice jean \
  --pocket_tts_device cpu

```

## Programmatic Usage in Python

Instantiate the pipeline directly in Python to specify pre-trained model identifiers:

```python
from speech_to_speech import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
    tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    llm_model_name="gpt-5.4-mini",
)

pipeline.run()

```

The constructor forwards these identifiers to the same resolution logic used by the CLI, automatically pulling assets from the Hugging Face Hub.

## Key Source Files Containing Model Logic

The following files contain the logic for selecting, resolving, and loading pre-trained models:

- [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) – Core TTS handler with MLX quantization resolution
- [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py) – Kokoro-82M backend implementation
- [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) – Pocket TTS backend with voice cloning
- [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) – ChatTTS backend for English/Chinese
- [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) – MMS TTS backend for multilingual synthesis
- [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) – Platform-specific STT model selection
- [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) – OpenAI-compatible LLM wrapper
- [`src/speech_to_speech/VAD/silero_vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/silero_vad_handler.py) – Automatic VAD model loading

## Summary

- Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and referenced by repository ID (e.g., `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`, `nvidia/parakeet-tdt-0.6b-v3`).
- The `huggingface/speech-to-speech` pipeline automatically downloads and caches these models on first use.
- Default models cover the full stack: Silero VAD (VAD), Parakeet TDT (STT), GPT-5.4-mini (LLM), and Qwen3-TTS (TTS).
- Alternative TTS backends (Kokoro, Pocket, ChatTTS, MMS) are available via CLI flags and handled by dedicated files in `src/speech_to_speech/TTS/`.
- Platform-specific resolution (MLX vs. PyTorch) occurs automatically in handlers like [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py) and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py).

## Frequently Asked Questions

### Where are the pre-trained model weights stored locally?

The pipeline uses the Hugging Face Hub cache directory (typically `~/.cache/huggingface/hub/` on Linux/macOS or `%USERPROFILE%\.cache\huggingface\hub` on Windows). When you reference a model by name, the library downloads the weights to this location and loads them from local storage on subsequent runs.

### Can I use custom models not listed in the defaults?

Yes. Any compatible model hosted on the Hugging Face Hub can be referenced using the `--*_model_name` flags (e.g., `--stt_model_name`, `--qwen3_tts_model_name`). The handler will attempt to load the specified repository ID, provided the model architecture matches the handler's expectations (e.g., Parakeet architecture for the STT handler).

### How does the pipeline choose between MLX and standard PyTorch models?

The resolution logic is platform-dependent. For example, in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) (lines 144-152), the code checks the operating system and defaults to `mlx-community/parakeet-tdt-0.6b-v3` on macOS and `nvidia/parakeet-tdt-0.6b-v3` on Linux/Windows. Similarly, the Qwen3-TTS handler applies MLX-specific quantization when running on Apple Silicon.

### Do I need an API key for all pre-trained models?

No. Only the OpenAI LLM component requires an API key (`OPENAI_API_KEY`). The VAD, STT, and TTS models (including Silero, Parakeet, Qwen3-TTS, Kokoro, and Pocket) are downloaded directly from the Hugging Face Hub and run locally without external API calls.