Pre-Trained Models for Speech-to-Speech: Complete Guide to the Hugging Face Pipeline

The huggingface/speech-to-speech pipeline ships with ready-to-use pre-trained models for every component: Silero VAD for voice detection, Parakeet TDT for STT, OpenAI GPT for LLM reasoning, and Qwen3-TTS as the default text-to-speech engine—with four additional optional TTS backends available.

The huggingface/speech-to-speech repository provides a complete, modular voice-agent stack with curated pre-trained models hosted on the Hugging Face Hub. Each component—VAD, STT, LLM, and TTS—has a battle-tested default that downloads automatically on first use, while alternative backends can be swapped via CLI flags or Python parameters.

Default Pre-Trained Models by Component

The pipeline uses these identifiers as factory defaults, resolving platform-specific variants where needed:

Voice Activity Detection (VAD)

Speech-to-Text (STT)

  • Parakeet TDT 0.6B v3

    The src/speech_to_speech/STT/parakeet_tdt_handler.py handler implements platform-aware resolution: on macOS it selects mlx-community/parakeet-tdt-0.6b-v3 for MLX acceleration; otherwise it falls back to nvidia/parakeet-tdt-0.6b-v3 for standard PyTorch inference.

Large Language Model (LLM)

Text-to-Speech (TTS) — Default

  • Qwen3-TTS 12Hz 1.7B CustomVoice (Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice)

    Implemented in src/speech_to_speech/TTS/qwen3_tts_handler.py. The _resolve_mlx_model_name function (lines 286-293) automatically appends a -6bit quantization suffix if the provided model name lacks one, optimizing for local inference.

Optional Pre-Trained TTS Backends

The pipeline supports four additional TTS engines, each with dedicated handlers:

Backend Hub Identifier Handler Location
Kokoro-82M hexgrad/Kokoro-82M src/speech_to_speech/TTS/kokoro_handler.py
Pocket TTS kyutai-labs/pocket-tts src/speech_to_speech/TTS/pocket_tts_handler.py
ChatTTS 2noise/ChatTTS src/speech_to_speech/TTS/chatTTS_handler.py
MMS TTS facebook/mms-tts src/speech_to_speech/TTS/facebookmms_handler.py

How Model Selection Works in the Source Code

The pipeline's model resolution logic determines which weights to fetch based on runtime context:

  1. CLI flags override defaults: Arguments like --stt_model_name, --qwen3_tts_model_name, or --kokoro_model_name are parsed and passed to handler constructors.

  2. Platform detection: As seen in src/speech_to_speech/STT/parakeet_tdt_handler.py (lines 144-152), the code checks sys.platform to choose between MLX-optimized and standard PyTorch variants.

  3. Automatic quantization: The Qwen3-TTS handler inspects model names and modifies them for efficient local execution without manual user intervention.

CLI Examples with Pre-Trained Models

Run with Default Models

export OPENAI_API_KEY=YOUR_KEY
speech-to-speech

Starts the WebSocket server with Parakeet TDT + Qwen3-TTS + OpenAI LLM, downloading any missing weights automatically.

Switch to Kokoro TTS

speech-to-speech \
  --tts kokoro \
  --kokoro_model_name hexgrad/Kokoro-82M

The --tts kokoro flag routes to src/speech_to_speech/TTS/kokoro_handler.py, which loads the specified checkpoint (lines 104-108).

Use Pocket TTS with Voice Preset

speech-to-speech \
  --tts pocket \
  --pocket_tts_voice jean \
  --pocket_tts_device cpu

The pocket_tts_handler.py (lines 30-38) fetches kyutai-labs/pocket-tts and applies the selected voice configuration.

Specify Parakeet TDT Explicitly

speech-to-speech \
  --stt_model_name nvidia/parakeet-tdt-0.6b-v3

Bypasses platform auto-detection to force the NVIDIA-hosted variant.

Programmatic Model Selection

from speech_to_speech import SpeechToSpeechPipeline

pipeline = SpeechToSpeechPipeline(
    stt_model_name="nvidia/parakeet-tdt-0.6b-v3",
    tts_model_name="Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
    llm_model_name="gpt-5.4-mini",
)

pipeline.run()

The constructor forwards identifiers to the same resolution layer used by the CLI, ensuring consistent caching and download behavior.

Summary

  • Five TTS options: Qwen3-TTS (default), Kokoro-82M, Pocket TTS, ChatTTS, and MMS TTS
  • Single STT default: Parakeet TDT 0.6B v3 with automatic MLX/PyTorch selection
  • VAD: Silero VAD v5, loaded transparently
  • LLM: OpenAI GPT-5.4-mini via Responses API (API key required)
  • All models are Hugging Face Hub identifiers—no manual download needed
  • Platform optimization happens automatically in handler logic

Frequently Asked Questions

How do I list all available pre-trained models for speech-to-speech?

The huggingface/speech-to-speech repository does not ship with a discovery command. Refer to the handler source files—src/speech_to_speech/TTS/ contains five handlers with documented default IDs, and src/speech_to_speech/STT/ shows the Parakeet TDT variants. Any Hub-compatible identifier can be passed if the model follows the expected interface.

Can I use local models instead of downloading from the Hub?

Yes. Most handlers accept absolute paths in their *_model_name arguments. The Qwen3-TTS handler, for example, resolves local paths through the same _resolve_mlx_model_name logic, skipping remote download if the file exists locally.

Why does the STT model change on macOS?

The parakeet_tdt_handler.py contains platform-conditional code (lines 144-152) that selects mlx-community/parakeet-tdt-0.6b-v3 for Apple Silicon/macOS to leverage Apple's MLX framework for accelerated inference, while Linux and Windows use the NVIDIA-hosted PyTorch version.

Is there a fully open-source LLM option instead of OpenAI?

Not as a pre-configured default. The responses_api_language_model.py wrapper implements the OpenAI Responses API. To use open weights, you would need to implement a compatible handler following the same interface, or use the pipeline's modular structure to inject a local LLM via the language model abstraction layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →