# What Models Are Supported by Hugging Face Speech-to-Speech: Complete Model Guide

> Explore Hugging Face Speech-to-Speech models for voice detection, STT, LLMs, and TTS. Discover supported options like Whisper, OpenAI APIs, Qwen3-TTS, and Kokoro.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: api-reference
- Published: 2026-08-01

---

**Hugging Face Speech-to-Speech supports modular backends across four pipeline stages: Silero VAD for voice detection, six STT options including Parakeet TDT and Whisper variants, OpenAI-compatible APIs and Transformers for LLMs, and five TTS models including Qwen3-TTS and Kokoro.**

The Hugging Face Speech-to-Speech (S2S) repository is a fully modular pipeline that lets you swap any of its four core stages—Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS). Understanding what models are supported by Hugging Face Speech-to-Speech is essential for building low-latency voice applications on CUDA, CPU, or Apple Silicon.

## Voice Activity Detection (VAD) Models

The pipeline uses **Silero VAD v5** as its sole built-in voice activity detection backend. According to the source code in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) [lines 53-60], the handler wraps the Silero model and feeds start/stop timestamps to the pipeline queue. This backend runs on all platforms without additional dependencies.

## Speech-to-Text (STT) Models

The STT layer supports six distinct backends, each implemented as a separate handler class inheriting from `BaseSTTHandler` in [`src/speech_to_speech/STT/base_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/base_stt_handler.py).

### Parakeet TDT (Default)

**Parakeet TDT** serves as the default STT backend, optimized for both CUDA/CPU via nano-parakeet and Apple Silicon via MLX. The implementation lives in [`src/speech_to_speech/STT/parakeet_tdt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/parakeet_tdt_handler.py) [lines 91-100], which handles audio chunking and model inference.

### Whisper Variants

The pipeline supports multiple Whisper implementations:

- **Whisper (Transformers)**: The standard implementation in [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) [lines 35-42] runs on CUDA/CPU and ships built-in.
- **Faster Whisper**: An optimized version using the `faster-whisper` library, implemented in [`src/speech_to_speech/STT/faster_whisper_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/faster_whisper_handler.py) [lines 19-26]. Install via the `faster-whisper` extra.
- **Lightning Whisper MLX**: Apple Silicon optimization found in [`src/speech_to_speech/STT/lightning_whisper_mlx_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/lightning_whisper_mlx_handler.py) [lines 36-43]. Install via the `whisper-mlx` extra.
- **MLX Audio Whisper**: Native macOS implementation that ships built-in on Apple Silicon.

### Paraformer

**Paraformer** via FunASR supports CUDA/CPU and excels at Chinese speech recognition. The handler in [`src/speech_to_speech/STT/paraformer_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/paraformer_handler.py) [lines 22-30] requires the `paraformer` extra for installation.

## Language Model (LLM) Backends

The LLM stage supports both API-based and local inference through handlers defined in [`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py) [lines 145-150].

### OpenAI-Compatible APIs

The **Responses API** and **Chat Completions** handlers—[`src/speech_to_speech/LLM/responses_api_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/responses_api_language_model.py) and [`src/speech_to_speech/LLM/chat_completions_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/chat_completions_language_model.py)—support any OpenAI-compatible endpoint including OpenAI, Hugging Face Inference Providers, OpenRouter, vLLM, and llama.cpp. These ship built-in and require only an API key or base URL.

### Local Transformers and MLX

- **Transformers**: Any text-generation model from the Hugging Face Hub runs locally via the built-in handler.
- **mlx-lm**: Optimized Apple Silicon inference that ships built-in on macOS.

## Text-to-Speech (TTS) Models

The TTS layer offers five backends, each implementing `BaseHandler[TTSIn, TTSOut]`.

### Qwen3-TTS (Default)

**Qwen3-TTS** serves as the default backend, automatically selecting **GGML** on Linux or **mlx-audio** on macOS. The handler in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) [lines 80-88] streams audio chunks back to the pipeline. Configuration arguments reside in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py) [lines 5-12].

### Kokoro-82M

**Kokoro-82M** runs on CUDA/CPU and Apple Silicon. The handler in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py) [lines 76-84] requires the `kokoro` extra on non-macOS systems but ships built-in on macOS.

### Optional TTS Backends

- **Pocket TTS**: CPU/CUDA streaming synthesis via [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) [lines 21-29]. Install with the `pocket` extra.
- **ChatTTS**: Conversational TTS implemented in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) [lines 28-36]. Install with the `chattts` extra.
- **MMS TTS**: Meta Music Speech synthesis in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py) [lines 67-75]. Install with the `facebook-mms` extra.

## Configuration Examples

Below are practical CLI configurations demonstrating different model combinations.

### Default End-to-End Pipeline

Run the default stack (Parakeet TDT → OpenAI Responses API → Qwen3-TTS):

```bash
speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "gpt-4o-mini"

```

### Apple Silicon Local Deployment

Optimize for local Mac execution using MLX across all stages:

```bash
speech-to-speech \
    --local_mac_optimal_settings \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

```

This flag automatically selects Parakeet TDT, mlx-lm, and MLX Qwen3-TTS.

### Custom STT and TTS Combination

Use Paraformer for Chinese STT with local MLX LLM and Qwen3-TTS:

```bash
speech-to-speech \
    --stt paraformer \
    --stt_model_name nvidia/parakeet-tdt-0.6b-v3 \
    --llm_backend mlx-lm \
    --tts qwen3 \
    --language auto

```

### Self-Hosted LLM with vLLM

Point to a local vLLM instance:

```bash
speech-to-speech \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream

```

### Alternative TTS Backend

Configure Pocket TTS with specific voice and device settings:

```bash
speech-to-speech \
    --stt whisper-mlx \
    --llm_backend transformers \
    --tts pocket \
    --pocket_tts_voice jean \
    --pocket_tts_device cpu

```

## Summary

- **VAD**: Only Silero VAD v5 is supported, implemented in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py).
- **STT**: Six backends available including Parakeet TDT (default), four Whisper variants, and Paraformer, each with dedicated handlers in `src/speech_to_speech/STT/`.
- **LLM**: Supports OpenAI-compatible APIs, Transformers, and mlx-lm through modular handlers in `src/speech_to_speech/LLM/`.
- **TTS**: Five backends including Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS, implemented in `src/speech_to_speech/TTS/`.
- **Installation**: Most models ship built-in; optional extras include `faster-whisper`, `whisper-mlx`, `paraformer`, `kokoro`, `pocket`, `chattts`, and `facebook-mms`.

## Frequently Asked Questions

### What is the default model configuration for Hugging Face Speech-to-Speech?

The default configuration uses **Parakeet TDT** for STT, **OpenAI Responses API** for LLM inference, and **Qwen3-TTS** for speech synthesis. This setup requires no additional installation beyond the base package and automatically optimizes for your hardware (GGML on Linux, mlx-audio on macOS).

### Can I run Hugging Face Speech-to-Speech entirely with local models?

Yes. Use the `--local_mac_optimal_settings` flag on Apple Silicon to automatically select Parakeet TDT, mlx-lm, and MLX Qwen3-TTS. On CUDA/CPU systems, specify `--llm_backend transformers` with any Hugging Face text-generation model and `--stt parakeet-tdt` or `--stt whisper` for local STT.

### Which STT model works best on Apple Silicon?

For Apple Silicon, you have three optimized options: **Parakeet TDT** with MLX (built-in), **Lightning Whisper MLX** (install via `whisper-mlx` extra), or **MLX Audio Whisper** (built-in on macOS). Parakeet TDT offers the best balance of speed and accuracy for English, while Lightning Whisper MLX excels at multilingual tasks.

### How do I install optional models like Kokoro or ChatTTS?

Install optional TTS backends using pip extras: `pip install speech-to-speech[kokoro]` for Kokoro-82M, `pip install speech-to-speech[chattts]` for ChatTTS, and `pip install speech-to-speech[pocket]` for Pocket TTS. For STT extras, use `pip install speech-to-speech[faster-whisper]` or `pip install speech-to-speech[paraformer]`.