Where to Find Pre-Trained Speech-to-Speech Models: Hugging Face Hub Integration Guide

Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and can be referenced directly by repository ID in the CLI or Python API to automatically download and cache weights for VAD, STT, LLM, and TTS components.

The huggingface/speech-to-speech repository provides a production-ready pipeline that ships with pre-trained speech-to-speech models for every stage of voice processing. According to the source code, all default model weights are hosted on the Hugging Face Hub and resolve automatically when you reference them by name, triggering download and local caching on first use.

Default Pre-Trained Models Available Out-of-the-Box

The pipeline includes ready-to-use models for each component of the speech-to-speech stack.

Voice Activity Detection (VAD)

The pipeline uses Silero VAD v5 (snakers4/silero-vad) for voice activity detection. This model is loaded automatically by the VAD handler in src/speech_to_speech/VAD/silero_vad_handler.py without requiring manual configuration.

Speech-to-Text (STT)

For speech recognition, the default is Parakeet TDT 0.6B v3. The handler at src/speech_to_speech/STT/parakeet_tdt_handler.py implements platform-specific resolution: it selects mlx-community/parakeet-tdt-0.6b-v3 on macOS and nvidia/parakeet-tdt-0.6b-v3 on other platforms (lines 144-152).

Large Language Model (LLM)

The default language model is OpenAI GPT-5.4-mini (openai/gpt-5.4-mini), managed by src/speech_to_speech/LLM/responses_api_language_model.py. This component handles the text generation stage between speech recognition and synthesis.

Text-to-Speech (TTS)

The default TTS backend uses Qwen3-TTS 12Hz 1.7B CustomVoice (Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice). The handler in src/speech_to_speech/TTS/qwen3_tts_handler.py includes automatic quantization logic in the _resolve_mlx_model_name function (lines 286-293), which appends a -6bit suffix to the model name if not already specified.

Optional TTS Backends and Alternative Pre-Trained Models

The pipeline supports several alternative TTS models that can be swapped via CLI flags:

How Model Resolution Works in the Source Code

The repository implements automatic model resolution logic that determines whether to use MLX, GGML, or standard Transformers backends based on the current platform.

In src/speech_to_speech/TTS/qwen3_tts_handler.py, the _resolve_mlx_model_name function handles quantization defaults by adding a -6bit quantization suffix when the model name does not already specify one (lines 286-293). Similarly, the STT handler in src/speech_to_speech/STT/parakeet_tdt_handler.py selects between MLX and NVIDIA variants based on the operating system (lines 144-152).

When you reference a model by its Hugging Face Hub identifier, the library downloads the weights to the local cache on first use and loads them into the appropriate handler.

CLI Examples for Using Pre-Trained Models

You can override default models using the --*_model_name flags or switch backends with --tts, --stt, and --llm flags.

Run the Pipeline with Default Models

Start the WebSocket server using the default pre-trained models (Parakeet TDT + Qwen3-TTS + OpenAI LLM):

export OPENAI_API_KEY=YOUR_KEY
speech-to-speech

Use Kokoro TTS Instead of Qwen3

Switch to the Kokoro-82M model by selecting the Kokoro handler and specifying the repository ID:

speech-to-speech \
  --tts kokoro \
  --kokoro_model_name hexgrad/Kokoro-82M \
  --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

Configure Pocket TTS with a Specific Voice

Use the Pocket TTS backend with a voice preset and CPU inference:

speech-to-speech \
  --tts pocket \
  --pocket_tts_voice jean \
  --pocket_tts_device cpu

Programmatic Usage in Python

Instantiate the pipeline directly in Python to specify pre-trained model identifiers:

from speech_to_speech import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
    tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    llm_model_name="gpt-5.4-mini",
)

pipeline.run()

The constructor forwards these identifiers to the same resolution logic used by the CLI, automatically pulling assets from the Hugging Face Hub.

Key Source Files Containing Model Logic

The following files contain the logic for selecting, resolving, and loading pre-trained models:

Summary

  • Pre-trained speech-to-speech models are hosted on the Hugging Face Hub and referenced by repository ID (e.g., Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice, nvidia/parakeet-tdt-0.6b-v3).
  • The huggingface/speech-to-speech pipeline automatically downloads and caches these models on first use.
  • Default models cover the full stack: Silero VAD (VAD), Parakeet TDT (STT), GPT-5.4-mini (LLM), and Qwen3-TTS (TTS).
  • Alternative TTS backends (Kokoro, Pocket, ChatTTS, MMS) are available via CLI flags and handled by dedicated files in src/speech_to_speech/TTS/.
  • Platform-specific resolution (MLX vs. PyTorch) occurs automatically in handlers like parakeet_tdt_handler.py and qwen3_tts_handler.py.

Frequently Asked Questions

Where are the pre-trained model weights stored locally?

The pipeline uses the Hugging Face Hub cache directory (typically ~/.cache/huggingface/hub/ on Linux/macOS or %USERPROFILE%\.cache\huggingface\hub on Windows). When you reference a model by name, the library downloads the weights to this location and loads them from local storage on subsequent runs.

Can I use custom models not listed in the defaults?

Yes. Any compatible model hosted on the Hugging Face Hub can be referenced using the --*_model_name flags (e.g., --stt_model_name, --qwen3_tts_model_name). The handler will attempt to load the specified repository ID, provided the model architecture matches the handler's expectations (e.g., Parakeet architecture for the STT handler).

How does the pipeline choose between MLX and standard PyTorch models?

The resolution logic is platform-dependent. For example, in src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 144-152), the code checks the operating system and defaults to mlx-community/parakeet-tdt-0.6b-v3 on macOS and nvidia/parakeet-tdt-0.6b-v3 on Linux/Windows. Similarly, the Qwen3-TTS handler applies MLX-specific quantization when running on Apple Silicon.

Do I need an API key for all pre-trained models?

No. Only the OpenAI LLM component requires an API key (OPENAI_API_KEY). The VAD, STT, and TTS models (including Silero, Parakeet, Qwen3-TTS, Kokoro, and Pocket) are downloaded directly from the Hugging Face Hub and run locally without external API calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →