Pre-Trained Models for Speech-to-Speech: Complete Guide to the Hugging Face Pipeline
The huggingface/speech-to-speech pipeline ships with ready-to-use pre-trained models for every component: Silero VAD for voice detection, Parakeet TDT for STT, OpenAI GPT for LLM reasoning, and Qwen3-TTS as the default text-to-speech engine—with four additional optional TTS backends available.
The huggingface/speech-to-speech repository provides a complete, modular voice-agent stack with curated pre-trained models hosted on the Hugging Face Hub. Each component—VAD, STT, LLM, and TTS—has a battle-tested default that downloads automatically on first use, while alternative backends can be swapped via CLI flags or Python parameters.
Default Pre-Trained Models by Component
The pipeline uses these identifiers as factory defaults, resolving platform-specific variants where needed:
Voice Activity Detection (VAD)
-
Silero VAD v5 (
snakers4/silero-vad)Loaded automatically by
src/speech_to_speech/VAD/silero_vad_handler.py. No configuration required; the handler downloads and caches the ONNX weights on first detection pass.
Speech-to-Text (STT)
-
Parakeet TDT 0.6B v3
The
src/speech_to_speech/STT/parakeet_tdt_handler.pyhandler implements platform-aware resolution: on macOS it selectsmlx-community/parakeet-tdt-0.6b-v3for MLX acceleration; otherwise it falls back tonvidia/parakeet-tdt-0.6b-v3for standard PyTorch inference.
Large Language Model (LLM)
-
OpenAI GPT-5.4-mini (default)
Managed by
src/speech_to_speech/LLM/responses_api_language_model.py. RequiresOPENAI_API_KEYenvironment variable. The model identifiergpt-5.4-miniis passed directly to the Responses API.
Text-to-Speech (TTS) — Default
-
Qwen3-TTS 12Hz 1.7B CustomVoice (
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice)Implemented in
src/speech_to_speech/TTS/qwen3_tts_handler.py. The_resolve_mlx_model_namefunction (lines 286-293) automatically appends a-6bitquantization suffix if the provided model name lacks one, optimizing for local inference.
Optional Pre-Trained TTS Backends
The pipeline supports four additional TTS engines, each with dedicated handlers:
| Backend | Hub Identifier | Handler Location |
|---|---|---|
| Kokoro-82M | hexgrad/Kokoro-82M |
src/speech_to_speech/TTS/kokoro_handler.py |
| Pocket TTS | kyutai-labs/pocket-tts |
src/speech_to_speech/TTS/pocket_tts_handler.py |
| ChatTTS | 2noise/ChatTTS |
src/speech_to_speech/TTS/chatTTS_handler.py |
| MMS TTS | facebook/mms-tts |
src/speech_to_speech/TTS/facebookmms_handler.py |
How Model Selection Works in the Source Code
The pipeline's model resolution logic determines which weights to fetch based on runtime context:
-
CLI flags override defaults: Arguments like
--stt_model_name,--qwen3_tts_model_name, or--kokoro_model_nameare parsed and passed to handler constructors. -
Platform detection: As seen in
src/speech_to_speech/STT/parakeet_tdt_handler.py(lines 144-152), the code checkssys.platformto choose between MLX-optimized and standard PyTorch variants. -
Automatic quantization: The Qwen3-TTS handler inspects model names and modifies them for efficient local execution without manual user intervention.
CLI Examples with Pre-Trained Models
Run with Default Models
export OPENAI_API_KEY=YOUR_KEY
speech-to-speech
Starts the WebSocket server with Parakeet TDT + Qwen3-TTS + OpenAI LLM, downloading any missing weights automatically.
Switch to Kokoro TTS
speech-to-speech \
--tts kokoro \
--kokoro_model_name hexgrad/Kokoro-82M
The --tts kokoro flag routes to src/speech_to_speech/TTS/kokoro_handler.py, which loads the specified checkpoint (lines 104-108).
Use Pocket TTS with Voice Preset
speech-to-speech \
--tts pocket \
--pocket_tts_voice jean \
--pocket_tts_device cpu
The pocket_tts_handler.py (lines 30-38) fetches kyutai-labs/pocket-tts and applies the selected voice configuration.
Specify Parakeet TDT Explicitly
speech-to-speech \
--stt_model_name nvidia/parakeet-tdt-0.6b-v3
Bypasses platform auto-detection to force the NVIDIA-hosted variant.
Programmatic Model Selection
from speech_to_speech import SpeechToSpeechPipeline
pipeline = SpeechToSpeechPipeline(
stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
llm_model_name="gpt-5.4-mini",
)
pipeline.run()
The constructor forwards identifiers to the same resolution layer used by the CLI, ensuring consistent caching and download behavior.
Summary
- Five TTS options: Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS
- Single STT default: Parakeet TDT 0.6B v3 with automatic MLX/PyTorch selection
- VAD: Silero VAD v5, loaded transparently
- LLM: OpenAI GPT-5.4-mini via Responses API (API key required)
- All models are Hugging Face Hub identifiers—no manual download needed
- Platform optimization happens automatically in handler logic
Frequently Asked Questions
How do I list all available pre-trained models for speech-to-speech?
The huggingface/speech-to-speech repository does not ship with a discovery command. Refer to the handler source files—src/speech_to_speech/TTS/ contains five handlers with documented default IDs, and src/speech_to_speech/STT/ shows the Parakeet TDT variants. Any Hub-compatible identifier can be passed if the model follows the expected interface.
Can I use local models instead of downloading from the Hub?
Yes. Most handlers accept absolute paths in their *_model_name arguments. The Qwen3-TTS handler, for example, resolves local paths through the same _resolve_mlx_model_name logic, skipping remote download if the file exists locally.
Why does the STT model change on macOS?
The parakeet_tdt_handler.py contains platform-conditional code (lines 144-152) that selects mlx-community/parakeet-tdt-0.6b-v3 for Apple Silicon/macOS to leverage Apple's MLX framework for accelerated inference, while Linux and Windows use the NVIDIA-hosted PyTorch version.
Is there a fully open-source LLM option instead of OpenAI?
Not as a pre-configured default. The responses_api_language_model.py wrapper implements the OpenAI Responses API. To use open weights, you would need to implement a compatible handler following the same interface, or use the pipeline's modular structure to inject a local LLM via the language model abstraction layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →