Where to Find Pre-Trained Speech-to-Speech Models: Hugging Face Hub Integration Guide
Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and can be referenced directly by repository ID in the CLI or Python API to automatically download and cache weights for VAD, STT, LLM, and TTS components.
The huggingface/speech-to-speech repository provides a production-ready pipeline that ships with pre-trained speech-to-speech models for every stage of voice processing. According to the source code, all default model weights are hosted on the Hugging Face Hub and resolve automatically when you reference them by name, triggering download and local caching on first use.
Default Pre-Trained Models Available Out-of-the-Box
The pipeline includes ready-to-use models for each component of the speech-to-speech stack.
Voice Activity Detection (VAD)
The pipeline uses Silero VAD v5 (snakers4/silero-vad) for voice activity detection. This model is loaded automatically by the VAD handler in src/speech_to_speech/VAD/silero_vad_handler.py without requiring manual configuration.
Speech-to-Text (STT)
For speech recognition, the default is Parakeet TDT 0.6B v3. The handler at src/speech_to_speech/STT/parakeet_tdt_handler.py implements platform-specific resolution: it selects mlx-community/parakeet-tdt-0.6b-v3 on macOS and nvidia/parakeet-tdt-0.6b-v3 on other platforms (lines 144-152).
Large Language Model (LLM)
The default language model is OpenAI GPT-5.4-mini (openai/gpt-5.4-mini), managed by src/speech_to_speech/LLM/responses_api_language_model.py. This component handles the text generation stage between speech recognition and synthesis.
Text-to-Speech (TTS)
The default TTS backend uses Qwen3-TTS 12Hz 1.7B CustomVoice (Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice). The handler in src/speech_to_speech/TTS/qwen3_tts_handler.py includes automatic quantization logic in the _resolve_mlx_model_name function (lines 286-293), which appends a -6bit suffix to the model name if not already specified.
Optional TTS Backends and Alternative Pre-Trained Models
The pipeline supports several alternative TTS models that can be swapped via CLI flags:
- Kokoro-82M (
hexgrad/Kokoro-82M) – Implemented insrc/speech_to_speech/TTS/kokoro_handler.py - Pocket TTS (
kyutai-labs/pocket-tts) – Implemented insrc/speech_to_speech/TTS/pocket_tts_handler.pywith voice cloning capabilities - ChatTTS (
2noise/ChatTTS) – Implemented insrc/speech_to_speech/TTS/chatTTS_handler.pyfor English and Chinese synthesis - MMS TTS (
facebook/mms-tts) – Implemented insrc/speech_to_speech/TTS/facebookmms_handler.pyfor multilingual support
How Model Resolution Works in the Source Code
The repository implements automatic model resolution logic that determines whether to use MLX, GGML, or standard Transformers backends based on the current platform.
In src/speech_to_speech/TTS/qwen3_tts_handler.py, the _resolve_mlx_model_name function handles quantization defaults by adding a -6bit quantization suffix when the model name does not already specify one (lines 286-293). Similarly, the STT handler in src/speech_to_speech/STT/parakeet_tdt_handler.py selects between MLX and NVIDIA variants based on the operating system (lines 144-152).
When you reference a model by its Hugging Face Hub identifier, the library downloads the weights to the local cache on first use and loads them into the appropriate handler.
CLI Examples for Using Pre-Trained Models
You can override default models using the --*_model_name flags or switch backends with --tts, --stt, and --llm flags.
Run the Pipeline with Default Models
Start the WebSocket server using the default pre-trained models (Parakeet TDT + Qwen3-TTS + OpenAI LLM):
export OPENAI_API_KEY=YOUR_KEY
speech-to-speech
Use Kokoro TTS Instead of Qwen3
Switch to the Kokoro-82M model by selecting the Kokoro handler and specifying the repository ID:
speech-to-speech \
--tts kokoro \
--kokoro_model_name hexgrad/Kokoro-82M \
--qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Configure Pocket TTS with a Specific Voice
Use the Pocket TTS backend with a voice preset and CPU inference:
speech-to-speech \
--tts pocket \
--pocket_tts_voice jean \
--pocket_tts_device cpu
Programmatic Usage in Python
Instantiate the pipeline directly in Python to specify pre-trained model identifiers:
from speech_to_speech import SpeechToSpeechPipeline
pipeline = SpeechToSpeechPipeline(
stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
llm_model_name="gpt-5.4-mini",
)
pipeline.run()
The constructor forwards these identifiers to the same resolution logic used by the CLI, automatically pulling assets from the Hugging Face Hub.
Key Source Files Containing Model Logic
The following files contain the logic for selecting, resolving, and loading pre-trained models:
src/speech_to_speech/TTS/qwen3_tts_handler.py– Core TTS handler with MLX quantization resolutionsrc/speech_to_speech/TTS/kokoro_handler.py– Kokoro-82M backend implementationsrc/speech_to_speech/TTS/pocket_tts_handler.py– Pocket TTS backend with voice cloningsrc/speech_to_speech/TTS/chatTTS_handler.py– ChatTTS backend for English/Chinesesrc/speech_to_speech/TTS/facebookmms_handler.py– MMS TTS backend for multilingual synthesissrc/speech_to_speech/STT/parakeet_tdt_handler.py– Platform-specific STT model selectionsrc/speech_to_speech/LLM/responses_api_language_model.py– OpenAI-compatible LLM wrappersrc/speech_to_speech/VAD/silero_vad_handler.py– Automatic VAD model loading
Summary
- Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and referenced by repository ID (e.g.,
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice,nvidia/parakeet-tdt-0.6b-v3). - The
huggingface/speech-to-speechpipeline automatically downloads and caches these models on first use. - Default models cover the full stack: Silero VAD (VAD), Parakeet TDT (STT), GPT-5.4-mini (LLM), and Qwen3-TTS (TTS).
- Alternative TTS backends (Kokoro, Pocket, ChatTTS, MMS) are available via CLI flags and handled by dedicated files in
src/speech_to_speech/TTS/. - Platform-specific resolution (MLX vs. PyTorch) occurs automatically in handlers like
parakeet_tdt_handler.pyandqwen3_tts_handler.py.
Frequently Asked Questions
Where are the pre-trained model weights stored locally?
The pipeline uses the Hugging Face Hub cache directory (typically ~/.cache/huggingface/hub/ on Linux/macOS or %USERPROFILE%\.cache\huggingface\hub on Windows). When you reference a model by name, the library downloads the weights to this location and loads them from local storage on subsequent runs.
Can I use custom models not listed in the defaults?
Yes. Any compatible model hosted on the Hugging Face Hub can be referenced using the --*_model_name flags (e.g., --stt_model_name, --qwen3_tts_model_name). The handler will attempt to load the specified repository ID, provided the model architecture matches the handler's expectations (e.g., Parakeet architecture for the STT handler).
How does the pipeline choose between MLX and standard PyTorch models?
The resolution logic is platform-dependent. For example, in src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 144-152), the code checks the operating system and defaults to mlx-community/parakeet-tdt-0.6b-v3 on macOS and nvidia/parakeet-tdt-0.6b-v3 on Linux/Windows. Similarly, the Qwen3-TTS handler applies MLX-specific quantization when running on Apple Silicon.
Do I need an API key for all pre-trained models?
No. Only the OpenAI LLM component requires an API key (OPENAI_API_KEY). The VAD, STT, and TTS models (including Silero, Parakeet, Qwen3-TTS, Kokoro, and Pocket) are downloaded directly from the Hugging Face Hub and run locally without external API calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →