What Models Are Supported by Hugging Face Speech-to-Speech: Complete Model Guide
Hugging Face Speech-to-Speech supports modular backends across four pipeline stages: Silero VAD for voice detection, six STT options including Parakeet TDT and Whisper variants, OpenAI-compatible APIs and Transformers for LLMs, and five TTS models including Qwen3-TTS and Kokoro.
The Hugging Face Speech-to-Speech (S2S) repository is a fully modular pipeline that lets you swap any of its four core stages—Voice Activity Detection (VAD), Speech-to-Text (STT), Language Model (LLM), and Text-to-Speech (TTS). Understanding what models are supported by Hugging Face Speech-to-Speech is essential for building low-latency voice applications on CUDA, CPU, or Apple Silicon.
Voice Activity Detection (VAD) Models
The pipeline uses Silero VAD v5 as its sole built-in voice activity detection backend. According to the source code in src/speech_to_speech/VAD/vad_handler.py [lines 53-60], the handler wraps the Silero model and feeds start/stop timestamps to the pipeline queue. This backend runs on all platforms without additional dependencies.
Speech-to-Text (STT) Models
The STT layer supports six distinct backends, each implemented as a separate handler class inheriting from BaseSTTHandler in src/speech_to_speech/STT/base_stt_handler.py.
Parakeet TDT (Default)
Parakeet TDT serves as the default STT backend, optimized for both CUDA/CPU via nano-parakeet and Apple Silicon via MLX. The implementation lives in src/speech_to_speech/STT/parakeet_tdt_handler.py [lines 91-100], which handles audio chunking and model inference.
Whisper Variants
The pipeline supports multiple Whisper implementations:
- Whisper (Transformers): The standard implementation in
src/speech_to_speech/STT/whisper_stt_handler.py[lines 35-42] runs on CUDA/CPU and ships built-in. - Faster Whisper: An optimized version using the
faster-whisperlibrary, implemented insrc/speech_to_speech/STT/faster_whisper_handler.py[lines 19-26]. Install via thefaster-whisperextra. - Lightning Whisper MLX: Apple Silicon optimization found in
src/speech_to_speech/STT/lightning_whisper_mlx_handler.py[lines 36-43]. Install via thewhisper-mlxextra. - MLX Audio Whisper: Native macOS implementation that ships built-in on Apple Silicon.
Paraformer
Paraformer via FunASR supports CUDA/CPU and excels at Chinese speech recognition. The handler in src/speech_to_speech/STT/paraformer_handler.py [lines 22-30] requires the paraformer extra for installation.
Language Model (LLM) Backends
The LLM stage supports both API-based and local inference through handlers defined in src/speech_to_speech/LLM/language_model.py [lines 145-150].
OpenAI-Compatible APIs
The Responses API and Chat Completions handlers—src/speech_to_speech/LLM/responses_api_language_model.py and src/speech_to_speech/LLM/chat_completions_language_model.py—support any OpenAI-compatible endpoint including OpenAI, Hugging Face Inference Providers, OpenRouter, vLLM, and llama.cpp. These ship built-in and require only an API key or base URL.
Local Transformers and MLX
- Transformers: Any text-generation model from the Hugging Face Hub runs locally via the built-in handler.
- mlx-lm: Optimized Apple Silicon inference that ships built-in on macOS.
Text-to-Speech (TTS) Models
The TTS layer offers five backends, each implementing BaseHandler[TTSIn, TTSOut].
Qwen3-TTS (Default)
Qwen3-TTS serves as the default backend, automatically selecting GGML on Linux or mlx-audio on macOS. The handler in src/speech_to_speech/TTS/qwen3_tts_handler.py [lines 80-88] streams audio chunks back to the pipeline. Configuration arguments reside in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py [lines 5-12].
Kokoro-82M
Kokoro-82M runs on CUDA/CPU and Apple Silicon. The handler in src/speech_to_speech/TTS/kokoro_handler.py [lines 76-84] requires the kokoro extra on non-macOS systems but ships built-in on macOS.
Optional TTS Backends
- Pocket TTS: CPU/CUDA streaming synthesis via
src/speech_to_speech/TTS/pocket_tts_handler.py[lines 21-29]. Install with thepocketextra. - ChatTTS: Conversational TTS implemented in
src/speech_to_speech/TTS/chatTTS_handler.py[lines 28-36]. Install with thechatttsextra. - MMS TTS: Meta Music Speech synthesis in
src/speech_to_speech/TTS/facebookmms_handler.py[lines 67-75]. Install with thefacebook-mmsextra.
Configuration Examples
Below are practical CLI configurations demonstrating different model combinations.
Default End-to-End Pipeline
Run the default stack (Parakeet TDT → OpenAI Responses API → Qwen3-TTS):
speech-to-speech \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--model_name "gpt-4o-mini"
Apple Silicon Local Deployment
Optimize for local Mac execution using MLX across all stages:
speech-to-speech \
--local_mac_optimal_settings \
--model_name mlx-community/Qwen3-4B-Instruct-2507-bf16
This flag automatically selects Parakeet TDT, mlx-lm, and MLX Qwen3-TTS.
Custom STT and TTS Combination
Use Paraformer for Chinese STT with local MLX LLM and Qwen3-TTS:
speech-to-speech \
--stt paraformer \
--stt_model_name nvidia/parakeet-tdt-0.6b-v3 \
--llm_backend mlx-lm \
--tts qwen3 \
--language auto
Self-Hosted LLM with vLLM
Point to a local vLLM instance:
speech-to-speech \
--stt parakeet-tdt \
--llm_backend chat-completions \
--model_name "Qwen/Qwen3-4B-Instruct-2507" \
--responses_api_base_url "http://localhost:8000/v1" \
--responses_api_stream
Alternative TTS Backend
Configure Pocket TTS with specific voice and device settings:
speech-to-speech \
--stt whisper-mlx \
--llm_backend transformers \
--tts pocket \
--pocket_tts_voice jean \
--pocket_tts_device cpu
Summary
- VAD: Only Silero VAD v5 is supported, implemented in
src/speech_to_speech/VAD/vad_handler.py. - STT: Six backends available including Parakeet TDT (default), four Whisper variants, and Paraformer, each with dedicated handlers in
src/speech_to_speech/STT/. - LLM: Supports OpenAI-compatible APIs, Transformers, and mlx-lm through modular handlers in
src/speech_to_speech/LLM/. - TTS: Five backends including Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS, implemented in
src/speech_to_speech/TTS/. - Installation: Most models ship built-in; optional extras include
faster-whisper,whisper-mlx,paraformer,kokoro,pocket,chattts, andfacebook-mms.
Frequently Asked Questions
What is the default model configuration for Hugging Face Speech-to-Speech?
The default configuration uses Parakeet TDT for STT, OpenAI Responses API for LLM inference, and Qwen3-TTS for speech synthesis. This setup requires no additional installation beyond the base package and automatically optimizes for your hardware (GGML on Linux, mlx-audio on macOS).
Can I run Hugging Face Speech-to-Speech entirely with local models?
Yes. Use the --local_mac_optimal_settings flag on Apple Silicon to automatically select Parakeet TDT, mlx-lm, and MLX Qwen3-TTS. On CUDA/CPU systems, specify --llm_backend transformers with any Hugging Face text-generation model and --stt parakeet-tdt or --stt whisper for local STT.
Which STT model works best on Apple Silicon?
For Apple Silicon, you have three optimized options: Parakeet TDT with MLX (built-in), Lightning Whisper MLX (install via whisper-mlx extra), or MLX Audio Whisper (built-in on macOS). Parakeet TDT offers the best balance of speed and accuracy for English, while Lightning Whisper MLX excels at multilingual tasks.
How do I install optional models like Kokoro or ChatTTS?
Install optional TTS backends using pip extras: pip install speech-to-speech[kokoro] for Kokoro-82M, pip install speech-to-speech[chattts] for ChatTTS, and pip install speech-to-speech[pocket] for Pocket TTS. For STT extras, use pip install speech-to-speech[faster-whisper] or pip install speech-to-speech[paraformer].
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →