How to Perform Voice Conversion with HuggingFace Speech-to-Speech: A Complete Guide
Voice conversion in HuggingFace speech-to-speech is accomplished by cloning a reference audio file or using a pre-computed speaker embedding as the target speaker for the Qwen3-TTS stage.
This open-source pipeline enables real-time voice transformation by extracting voice features from a reference recording and applying them to synthesized speech. Whether you want to mimic a specific speaker or design a custom voice, the huggingface/speech-to-speech repository provides a streamlined interface through command-line arguments and modular handlers.
Setting Up Voice Conversion Arguments
Voice conversion behavior is controlled through dedicated CLI flags defined in src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py. The key parameters for voice cloning are:
--qwen3_tts_ref_audio– Path to a reference audio file (WAV format)--qwen3_tts_ref_spk– Path to a pre-computed speaker embedding (.spkfile)--qwen3_tts_ref_rvq– Path to pre-computed RVQ tokens (.rvqfile)--qwen3_tts_ref_cache_dir– Directory to cache speaker embeddings for faster subsequent runs--qwen3_tts_speaker– Built-in speaker preset (empty string forces cloning mode)
These arguments are parsed into Qwen3TTSHandlerArguments and passed through the pipeline initialization in src/speech_to_speech/s2s_pipeline.py【/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py#L61-L78】【/src/speech_to_speech/s2s_pipeline.py#L24-L33】.
The Qwen3-TTS Handler: Core Voice Conversion Logic
The Qwen3TTSHandler class in src/speech_to_speech/TTS/qwen3_tts_handler.py implements the actual voice conversion. It normalizes reference paths, auto-detects the optimal backend (mlx-audio on Apple Silicon, faster-qwen3-tts elsewhere), and streams synthesized audio chunks【/src/speech_to_speech/TTS/qwen3_tts_handler.py#L94-L118】.
The handler supports three distinct generation modes:
- Voice cloning – Uses raw reference audio or cached embeddings to replicate a specific speaker's voice
- Custom voice – Selects from built-in speaker presets like
AidenorAda - Voice design – Allows instruction-guided voice creation via
--qwen3_tts_instruct
Backend-specific setup occurs in _setup_mlx() (Apple Silicon) or _setup_faster() (GGML/CUDA), where reference audio is processed into speaker embeddings.
Step-by-Step Voice Conversion Workflow
Step 1: Prepare Your Reference Audio
Record or obtain a clean audio sample of your target speaker. For best results, use:
- WAV format
- 5–30 seconds of clear speech
- Minimal background noise
Step 2: Run Basic Voice Conversion
Execute the pipeline with your reference file:
speech-to-speech \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3 \
--qwen3_tts_ref_audio /path/to/target_speaker.wav \
--qwen3_tts_speaker ""
Setting --qwen3_tts_speaker "" explicitly disables preset speakers and forces cloning mode.
Step 3: Cache Speaker Embeddings for Faster Startup
First run extracts and caches speaker features. Subsequent launches use pre-computed embeddings:
speech-to-speech \
--tts qwen3 \
--qwen3_tts_ref_audio /path/to/speaker.wav \
--qwen3_tts_ref_cache_dir ~/.cache/qwen3_voices
Step 4: Use Cached Embeddings Directly
Skip audio processing entirely by pointing to cached files:
speech-to-speech \
--tts qwen3 \
--qwen3_tts_ref_spk ~/.cache/qwen3_voices/speaker.spk \
--qwen3_tts_ref_rvq ~/.cache/qwen3_voices/speaker.rvq \
--qwen3_tts_ref_text "Hello, this is my cloned voice."
Platform-Specific Configuration
Apple Silicon (MLX Audio)
On macOS with Apple Silicon, the handler automatically selects mlx-audio with built-in quantization. Backend specification is ignored:
speech-to-speech \
--tts qwen3 \
--qwen3_tts_backend torch \
--qwen3_tts_ref_audio /Users/me/voice.wav
The torch backend flag has no effect on Darwin—mlx-audio is always used【/src/speech_to_speech/TTS/qwen3_tts_handler.py#L94-L118】.
Linux/CUDA (GGML/Faster Backend)
Other platforms use faster-qwen3-tts via the _setup_faster() initialization path, enabling GPU-accelerated voice conversion on NVIDIA hardware.
Real-Time Streaming Architecture
Voice conversion operates in streaming mode by default. The streaming_chunk_size parameter controls audio buffer sizes, enabling near-instantaneous playback as the LLM generates text and the TTS handler synthesizes speech. This architecture eliminates wait times for full utterance completion.
Summary
- Voice conversion in HuggingFace speech-to-speech clones a target speaker using reference audio or cached embeddings
- Core implementation resides in
Qwen3TTSHandler(src/speech_to_speech/TTS/qwen3_tts_handler.py) with backend auto-detection - CLI arguments in
qwen3_tts_arguments.pyexpose--qwen3_tts_ref_audio,--qwen3_tts_ref_spk, and caching options - Pipeline wiring occurs in
s2s_pipeline.py, connecting arguments to the active TTS handler - Three generation modes: voice cloning, custom presets, and instruction-based voice design
- Platform optimization:
mlx-audioon Apple Silicon,faster-qwen3-ttselsewhere
Frequently Asked Questions
What audio format works best for voice conversion reference files?
WAV format with 16-bit PCM encoding and 16kHz or 24kHz sample rate provides optimal results. The Qwen3-TTS handler performs internal resampling, but starting with clean, uncompressed audio minimizes artifacts. Five to thirty seconds of continuous speech captures sufficient prosodic characteristics.
How do I speed up voice conversion on repeated runs?
Use --qwen3_tts_ref_cache_dir to persist speaker embeddings (.spk) and RVQ tokens (.rvq). After the first conversion, reference the cached files directly with --qwen3_tts_ref_spk and --qwen3_tts_ref_rvq instead of reprocessing the original audio. This reduces initialization from seconds to milliseconds.
Can I convert voice without any reference audio?
Yes—use built-in speakers via --qwen3_tts_speaker <name> (e.g., Aiden, Ada) or supply an instruction prompt with --qwen3_tts_instruct for VoiceDesign models. However, true voice cloning requires either reference audio or pre-computed embeddings from a specific target speaker.
Does voice conversion work offline?
Yes, provided the Qwen3-TTS model weights are cached locally. The pipeline performs all inference on-device; no API calls are required after initial model download. The mlx-audio and faster-qwen3-tts backends both support fully offline operation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →