# Pre-Trained Models for Speech-to-Speech: Complete Guide to the Hugging Face Pipeline

> Explore pre-trained models for speech-to-speech with the Hugging Face pipeline. Discover Silero VAD, Parakeet TDT, OpenAI GPT, and Qwen3-TTS for your audio projects.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: tutorial
- Published: 2026-08-02

---

**The huggingface/speech-to-speech pipeline ships with ready-to-use pre-trained models for every component: Silero VAD for voice detection, Parakeet TDT for STT, OpenAI GPT for LLM reasoning, and Qwen3-TTS as the default text-to-speech engine—with four additional optional TTS backends available.**

The **huggingface/speech-to-speech** repository provides a complete, modular voice-agent stack with curated pre-trained models hosted on the Hugging Face Hub. Each component—VAD, STT, LLM, and TTS—has a battle-tested default that downloads automatically on first use, while alternative backends can be swapped via CLI flags or Python parameters.

## Default Pre-Trained Models by Component

The pipeline uses these identifiers as factory defaults, resolving platform-specific variants where needed:

### Voice Activity Detection (VAD)

- **Silero VAD v5** (`snakers4/silero-vad`)

  Loaded automatically by [`src/speech_to_speech/VAD/silero_vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/silero_vad_handler.py). No configuration required; the handler downloads and caches the ONNX weights on first detection pass.

### Speech-to-Text (STT)

- **Parakeet TDT 0.6B v3**

  The [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) handler implements platform-aware resolution: on macOS it selects `mlx-community/parakeet-tdt-0.6b-v3` for MLX acceleration; otherwise it falls back to `nvidia/parakeet-tdt-0.6b-v3` for standard PyTorch inference.

### Large Language Model (LLM)

- **OpenAI GPT-5.4-mini** (default)

  Managed by [`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py). Requires `OPENAI_API_KEY` environment variable. The model identifier `gpt-5.4-mini` is passed directly to the Responses API.

### Text-to-Speech (TTS) — Default

- **Qwen3-TTS 12Hz 1.7B CustomVoice** (`Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`)

  Implemented in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py). The `_resolve_mlx_model_name` function (lines 286-293) automatically appends a `-6bit` quantization suffix if the provided model name lacks one, optimizing for local inference.

## Optional Pre-Trained TTS Backends

The pipeline supports four additional TTS engines, each with dedicated handlers:

| Backend | Hub Identifier | Handler Location |
|---------|---------------|------------------|
| **Kokoro-82M** | `hexgrad/Kokoro-82M` | [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py) |
| **Pocket TTS** | `kyutai-labs/pocket-tts` | [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) |
| **ChatTTS** | `2noise/ChatTTS` | [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) |
| **MMS TTS** | `facebook/mms-tts` | [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) |

## How Model Selection Works in the Source Code

The pipeline's **model resolution logic** determines which weights to fetch based on runtime context:

1. **CLI flags override defaults**: Arguments like `--stt_model_name`, `--qwen3_tts_model_name`, or `--kokoro_model_name` are parsed and passed to handler constructors.

2. **Platform detection**: As seen in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) (lines 144-152), the code checks `sys.platform` to choose between MLX-optimized and standard PyTorch variants.

3. **Automatic quantization**: The Qwen3-TTS handler inspects model names and modifies them for efficient local execution without manual user intervention.

## CLI Examples with Pre-Trained Models

### Run with Default Models

```bash
export OPENAI_API_KEY=YOUR_KEY
speech-to-speech

```

Starts the WebSocket server with Parakeet TDT + Qwen3-TTS + OpenAI LLM, downloading any missing weights automatically.

### Switch to Kokoro TTS

```bash
speech-to-speech \
  --tts kokoro \
  --kokoro_model_name hexgrad/Kokoro-82M

```

The `--tts kokoro` flag routes to [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py), which loads the specified checkpoint (lines 104-108).

### Use Pocket TTS with Voice Preset

```bash
speech-to-speech \
  --tts pocket \
  --pocket_tts_voice jean \
  --pocket_tts_device cpu

```

The [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) (lines 30-38) fetches `kyutai-labs/pocket-tts` and applies the selected voice configuration.

### Specify Parakeet TDT Explicitly

```bash
speech-to-speech \
  --stt_model_name nvidia/parakeet-tdt-0.6b-v3

```

Bypasses platform auto-detection to force the NVIDIA-hosted variant.

## Programmatic Model Selection

```python
from speech_to_speech import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
    tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    llm_model_name="gpt-5.4-mini",
)

pipeline.run()

```

The constructor forwards identifiers to the same resolution layer used by the CLI, ensuring consistent caching and download behavior.

## Summary

- **Five TTS options**: Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS
- **Single STT default**: Parakeet TDT 0.6B v3 with automatic MLX/PyTorch selection
- **VAD**: Silero VAD v5, loaded transparently
- **LLM**: OpenAI GPT-5.4-mini via Responses API (API key required)
- **All models** are Hugging Face Hub identifiers—no manual download needed
- **Platform optimization** happens automatically in handler logic

## Frequently Asked Questions

### How do I list all available pre-trained models for speech-to-speech?

The huggingface/speech-to-speech repository does not ship with a discovery command. Refer to the handler source files—`src/speech_to_speech/TTS/` contains five handlers with documented default IDs, and `src/speech_to_speech/STT/` shows the Parakeet TDT variants. Any Hub-compatible identifier can be passed if the model follows the expected interface.

### Can I use local models instead of downloading from the Hub?

Yes. Most handlers accept absolute paths in their `*_model_name` arguments. The Qwen3-TTS handler, for example, resolves local paths through the same `_resolve_mlx_model_name` logic, skipping remote download if the file exists locally.

### Why does the STT model change on macOS?

The [`parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/parakeet_tdt_handler.py) contains platform-conditional code (lines 144-152) that selects `mlx-community/parakeet-tdt-0.6b-v3` for Apple Silicon/macOS to leverage Apple's MLX framework for accelerated inference, while Linux and Windows use the NVIDIA-hosted PyTorch version.

### Is there a fully open-source LLM option instead of OpenAI?

Not as a pre-configured default. The [`responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/responses_api_language_model.py) wrapper implements the OpenAI Responses API. To use open weights, you would need to implement a compatible handler following the same interface, or use the pipeline's modular structure to inject a local LLM via the language model abstraction layer.