Are There Pre-Trained Speech-to-Speech Models Available?

Yes, the huggingface/speech-to-speech library ships with ready-to-use wrappers for multiple pre-trained speech-to-speech models that download automatically from the Hugging Face Hub.

The huggingface/speech-to-speech repository provides a unified S2SPipeline that orchestrates end-to-end voice processing using pre-trained speech-to-speech models. You can initialize production-ready pipelines for text-to-speech, speech-to-text, and full speech-to-speech conversion without manual model downloads or complex configuration.

How the Pipeline Loads Pre-Trained Models

The architecture centers on the S2SPipeline class in src/speech_to_speech/s2s_pipeline.py, which acts as a high-level orchestrator. When you call from_pretrained(), the pipeline inspects the provided arguments, resolves platform-specific variants (such as Apple Silicon MLX quantization), and delegates to specialized handler classes that invoke from_pretrained() on the underlying model architectures.

Argument dataclasses expose model identifiers through CLI-friendly interfaces. For example, Qwen3TTSArguments in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py defines the qwen3_tts_model_name parameter with a default value of Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice.

Handler classes implement the actual loading logic. The Qwen3TTSHandler in src/speech_to_speech/TTS/qwen3_tts_handler.py supports both Faster-Transformer and MLX backends, automatically appending -6bit suffixes for MLX quantized variants when needed.

Supported Pre-Trained Model Families

The library supports diverse pre-trained speech models across the STT → LLM → TTS stack.

Qwen3-TTS Models

The primary TTS backend uses Qwen3-TTS models, loaded via Qwen3TTSHandler. The handler resolves model identifiers and initializes either FasterQwen3TTS for standard CUDA/CPU inference or MLX-optimized variants for Apple Silicon. Default models include Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice and MLX community variants like mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16.

Whisper STT Models

For speech-to-text, the pipeline integrates Whisper models through WhisperSTTArguments (defined in src/speech_to_speech/arguments_classes/whisper_stt_arguments.py) and WhisperSTTHandler (implemented in src/speech_to_speech/STT/whisper_stt_handler.py). The handler loads models using AutoProcessor.from_pretrained() and AutoModelForSpeechSeq2Seq.from_pretrained(), with defaults pointing to distil-whisper/distil-large-v3.

Alternative TTS Backends

Beyond Qwen3-TTS, the library exposes handlers for Parler-TTS, Parakeet, Paraformer, Facebook-MMS, Kokoro, Pocket TTS, and ChatTTS. Each backend provides its own arguments class (e.g., ParlerTTSArguments in src/speech_to_speech/arguments_classes/parler_tts_arguments.py) and handler implementation, following the same from_pretrained pattern.

Practical Code Examples

Initialize a pipeline with default pre-trained models:

from speech_to_speech.s2s_pipeline import S2SPipeline

pipeline = S2SPipeline.from_pretrained()
audio_output = pipeline.run_text("Hello, how are you?")

Specify a specific MLX-optimized model for Apple Silicon:

from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.qwen3_tts_arguments import Qwen3TTSArguments

args = Qwen3TTSArguments(
    qwen3_tts_model_name="mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16"
)
pipeline = S2SPipeline.from_pretrained(qwen3_tts_args=args)
audio_output = pipeline.run_text("Bonjour, je suis un modèle TTS.")

Use the Parler-TTS backend instead:

from speech_to_speech.s2s_pipeline import S2SPipeline
from speech_to_speech.arguments_classes.parler_tts_arguments import ParlerTTSArguments

args = ParlerTTSArguments(parler_tts_model_name="parler-tts/parler-mini-v1-jenny")
pipeline = S2SPipeline.from_pretrained(parler_tts_args=args)
audio_output = pipeline.run_text("Testing the Parler TTS model.")

All examples download model weights automatically on first execution.

Summary

  • The S2SPipeline in src/speech_to_speech/s2s_pipeline.py provides unified access to pre-trained speech-to-speech models.
  • Qwen3TTSHandler and WhisperSTTHandler manage model instantiation for TTS and STT components respectively.
  • Models download automatically from the Hugging Face Hub when calling from_pretrained().
  • Apple Silicon users receive optimized MLX variants through automatic suffix resolution (e.g., -6bit).

Frequently Asked Questions

What pre-trained models are included by default?

The pipeline defaults to Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice for text-to-speech and distil-whisper/distil-large-v3 for speech-to-text. These initialize automatically when you call S2SPipeline.from_pretrained() without arguments.

Can I use custom Hugging Face models?

Yes. Pass any valid Hugging Face Hub model identifier to the appropriate arguments class, such as Qwen3TTSArguments(qwen3_tts_model_name="your-username/your-model"). The handler validates the identifier and downloads weights via the standard from_pretrained() mechanism.

Does the library support quantized models for Apple Silicon?

Yes. When running on Apple Silicon, Qwen3TTSHandler detects the platform and automatically resolves MLX-quantized variants by appending -6bit to the model name, or you can explicitly specify MLX community models like mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16.

How do I switch between different TTS backends?

Import the specific arguments class for your desired backend (e.g., ParlerTTSArguments or ParakeetTDTArguments) and pass the instance to S2SPipeline.from_pretrained() using the appropriate parameter name (e.g., parler_tts_args or parakeet_tdt_args).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →