Speech Transformations in the Hugging Face Speech-to-Speech Pipeline
The Hugging Face Speech-to-Speech repository supports voice cloning, custom preset speakers, voice design from text descriptions, and multilingual TTS through modular handlers that transform input audio streams into new audio with different characteristics, languages, or styles.
The huggingface/speech-to-speech repository implements a modular real-time pipeline that converts input audio into transformed output audio. Built around interchangeable STT (speech-to-text), LLM, and TTS (text-to-speech) components, the system enables developers to perform voice conversion, language translation, and stylistic rewriting by routing audio through specialized handlers.
Core Speech Transformation Types
The pipeline supports several distinct speech transformation modes determined by the TTS handler configuration. These transformations modify voice characteristics while preserving or altering the spoken content.
Voice Cloning from Reference Audio
The voice cloning transformation reproduces the timbre of a reference speaker by extracting speaker embeddings from provided audio. This mode requires a reference audio file (ref_audio) and optional reference text (ref_text) to guide the cloning process.
In src/speech_to_speech/TTS/qwen3_tts_handler.py, the Qwen3TTSHandler._process_voice_clone method handles this transformation. The handler processes the reference audio to generate consistent speaker embeddings that synthesize speech matching the original voice characteristics.
args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"
Custom Voice Selection from Presets
The custom-voice transformation generates speech using predefined speakers from a catalog (e.g., "Aiden", "Maya"). This mode selects speaker embeddings from the model's internal library rather than extracting them from reference audio.
The Qwen3TTSHandler._process_custom_voice method in src/speech_to_speech/TTS/qwen3_tts_handler.py implements this through the _resolve_speaker helper function. Users specify the speaker via the qwen3_tts_speaker argument.
args.qwen3_tts_handler_kwargs.speaker = "Aiden"
args.module_kwargs.tts = "qwen3"
Voice Design via Text Instructions
The voice design transformation synthesizes novel voices from textual descriptions without requiring reference audio. Users provide an instruction prompt (instruct) describing desired characteristics such as "a calm, deep voice with slight echo."
The Qwen3TTSHandler._process_voice_design method processes these instructions to create new voice embeddings on-the-fly. This method is located at line 851 of src/speech_to_speech/TTS/qwen3_tts_handler.py.
args.qwen3_tts_handler_kwargs.instruct = "a calm, deep voice with slight echo"
args.module_kwargs.tts = "qwen3"
Generic and Specialized TTS Handlers
Beyond the Qwen3-TTS handler, the repository provides several specialized handlers for specific use cases:
- ChatTTS: Provides a default generic voice using randomly sampled speaker embeddings. Implemented in
src/speech_to_speech/TTS/chatTTS_handler.pyvia theChatTTSHandler.setupmethod. - Pocket TTS: Lightweight multilingual TTS available through the
speech-to-speech[TTS,pocket]optional dependency. Located insrc/speech_to_speech/TTS/pocket_tts_handler.py. - Kokoro TTS: Japanese-focused TTS with voice-style control via the
kokoro_tts_speakerparameter. Located insrc/speech_to_speech/TTS/kokoro_handler.py. - Facebook MMS TTS: Multilingual streaming TTS from Meta, available via
speech-to-speech[facebook-mms]. Located insrc/speech_to_speech/TTS/facebookmms_handler.py.
Pipeline Architecture for Advanced Transformations
The modular architecture enables complex speech transformations by combining TTS handlers with STT and LLM backends.
STT and LLM Integration
The pipeline supports multiple STT backends including Whisper, Faster-Whisper, Paraformer, MLX-Audio-Whisper, and Parakeet-TDT. These feed into LLM backends (Transformers, MLX-LM, OpenAI-compatible APIs) that can perform:
- Voice conversion: Preserve content while changing the speaker identity
- Language translation: STT → LLM translation → TTS in target language
- Stylistic rewriting: LLM modifies the transcript (e.g., "make it formal") before synthesis
- Realtime voice-over: Live microphone input processed through VAD → STT → LLM → TTS with speculative turn handling for low latency
All configurations are managed through argument classes in src/speech_to_speech/arguments_classes/, which define parameters for each handler type.
Configuration and Usage Examples
Complete Voice Cloning Setup
The following example demonstrates configuring the pipeline for voice cloning on Apple Silicon:
from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args, build_pipeline
from speech_to_speech.utils.thread_manager import ThreadManager
# Parse CLI args or construct them manually
args = parse_arguments()
# Configure for voice cloning
args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"
# Prepare arguments for all handlers
prepare_all_args(
args.module_kwargs,
args.whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
)
# Build and start pipeline
pipeline_manager: ThreadManager = build_pipeline(
args.module_kwargs,
args.socket_receiver_kwargs,
args.socket_sender_kwargs,
args.websocket_streamer_kwargs,
args.vad_handler_kwargs,
args.whisper_stt_handler_kwargs,
args.faster_whisper_stt_handler_kwargs,
args.paraformer_stt_handler_kwargs,
args.mlx_audio_whisper_stt_handler_kwargs,
args.parakeet_tdt_stt_handler_kwargs,
args.language_model_handler_kwargs,
args.responses_api_language_model_handler_kwargs,
args.chat_tts_handler_kwargs,
args.facebook_mms_tts_handler_kwargs,
args.pocket_tts_handler_kwargs,
args.kokoro_tts_handler_kwargs,
args.qwen3_tts_handler_kwargs,
initialize_queues_and_events(),
)
pipeline_manager.start()
Language Translation Pipeline
To combine speech transformation with translation:
# LLM configuration for translation
args.language_model_handler_kwargs.model_name = "facebook/opt-1.3b"
args.language_model_handler_kwargs.chat_prompt = "Translate English to French and respond."
# TTS configuration for French output
args.qwen3_tts_handler_kwargs.language = "fr"
args.module_kwargs.tts = "qwen3"
Quick-Start Generic TTS
For immediate deployment without reference audio:
args.module_kwargs.tts = "chatTTS"
This instantiates ChatTTSHandler, which automatically selects a random speaker embedding from the generic multilingual model.
Summary
- Voice cloning reproduces speaker timbre from reference audio using
Qwen3TTSHandler._process_voice_cloneinsrc/speech_to_speech/TTS/qwen3_tts_handler.py. - Custom voices select from preset speakers like "Aiden" via the
qwen3_tts_speakerparameter andQwen3TTSHandler._process_custom_voice. - Voice design creates novel voices from text descriptions using
Qwen3TTSHandler._process_voice_designand theinstructparameter. - Specialized handlers include ChatTTS, Pocket TTS, Kokoro TTS, and Facebook MMS TTS for specific languages and deployment constraints.
- Advanced transformations combine STT, LLM, and TTS components to enable real-time translation, voice conversion, and stylistic rewriting.
- All handlers expose a common
process()method returningint16PCM chunks, enabling seamless integration with output streamers likeLocalAudioStreamerorWebSocketStreamer.
Frequently Asked Questions
What audio format is required for voice cloning?
Voice cloning requires a reference audio file (ref_audio) in a standard format (typically WAV) and optionally reference text (ref_text) containing the transcript of the reference audio. The Qwen3TTSHandler extracts speaker embeddings from this audio to synthesize matching voice characteristics.
Can I use the speech-to-speech pipeline for real-time voice conversion?
Yes, the pipeline supports real-time voice-over by processing live microphone input through Voice Activity Detection (VAD) → STT → LLM → TTS with speculative turn handling for low latency. All TTS handlers return audio as streams of int16 PCM chunks suitable for real-time playback.
How do I switch between different TTS backends?
Set args.module_kwargs.tts to the desired handler name ("qwen3", "chatTTS", "pocket", "kokoro", or "facebook_mms") and ensure the corresponding handler arguments are prepared via prepare_all_args(). The pipeline automatically instantiates the correct handler class from src/speech_to_speech/TTS/ based on this configuration.
Does the pipeline support multilingual speech transformations?
Yes, the Qwen3-TTS handler supports multiple languages through the language parameter, while specialized handlers like Kokoro TTS (Japanese) and Facebook MMS TTS provide additional multilingual capabilities. When combined with LLM translation, the pipeline can perform cross-lingual voice conversion (e.g., English input to French output with cloned voice characteristics).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →