# How to Perform Voice Conversion with HuggingFace Speech-to-Speech: A Complete Guide

> Learn to perform voice conversion with HuggingFace speech-to-speech by cloning audio or using speaker embeddings for Qwen3-TTS. Get started with our complete guide today.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-08-02

---

**Voice conversion in HuggingFace speech-to-speech is accomplished by cloning a reference audio file or using a pre-computed speaker embedding as the target speaker for the Qwen3-TTS stage.**

This open-source pipeline enables real-time voice transformation by extracting voice features from a reference recording and applying them to synthesized speech. Whether you want to mimic a specific speaker or design a custom voice, the `huggingface/speech-to-speech` repository provides a streamlined interface through command-line arguments and modular handlers.

## Setting Up Voice Conversion Arguments

Voice conversion behavior is controlled through dedicated CLI flags defined in [`src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py). The key parameters for voice cloning are:

- `--qwen3_tts_ref_audio` – Path to a reference audio file (WAV format)
- `--qwen3_tts_ref_spk` – Path to a pre-computed speaker embedding (`.spk` file)
- `--qwen3_tts_ref_rvq` – Path to pre-computed RVQ tokens (`.rvq` file)
- `--qwen3_tts_ref_cache_dir` – Directory to cache speaker embeddings for faster subsequent runs
- `--qwen3_tts_speaker` – Built-in speaker preset (empty string forces cloning mode)

These arguments are parsed into `Qwen3TTSHandlerArguments` and passed through the pipeline initialization in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py)【/src/speech_to_speech/arguments_classes/qwen3_tts_arguments.py#L61-L78】【/src/speech_to_speech/s2s_pipeline.py#L24-L33】.

## The Qwen3-TTS Handler: Core Voice Conversion Logic

The `Qwen3TTSHandler` class in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) implements the actual voice conversion. It normalizes reference paths, auto-detects the optimal backend (`mlx-audio` on Apple Silicon, `faster-qwen3-tts` elsewhere), and streams synthesized audio chunks【/src/speech_to_speech/TTS/qwen3_tts_handler.py#L94-L118】.

The handler supports three distinct generation modes:

1. **Voice cloning** – Uses raw reference audio or cached embeddings to replicate a specific speaker's voice
2. **Custom voice** – Selects from built-in speaker presets like `Aiden` or `Ada`
3. **Voice design** – Allows instruction-guided voice creation via `--qwen3_tts_instruct`

Backend-specific setup occurs in `_setup_mlx()` (Apple Silicon) or `_setup_faster()` (GGML/CUDA), where reference audio is processed into speaker embeddings.

## Step-by-Step Voice Conversion Workflow

### Step 1: Prepare Your Reference Audio

Record or obtain a clean audio sample of your target speaker. For best results, use:
- WAV format
- 5–30 seconds of clear speech
- Minimal background noise

### Step 2: Run Basic Voice Conversion

Execute the pipeline with your reference file:

```bash
speech-to-speech \
  --stt parakeet-tdt \
  --llm_backend responses-api \
  --tts qwen3 \
  --qwen3_tts_ref_audio /path/to/target_speaker.wav \
  --qwen3_tts_speaker ""

```

Setting `--qwen3_tts_speaker ""` explicitly disables preset speakers and forces cloning mode.

### Step 3: Cache Speaker Embeddings for Faster Startup

First run extracts and caches speaker features. Subsequent launches use pre-computed embeddings:

```bash
speech-to-speech \
  --tts qwen3 \
  --qwen3_tts_ref_audio /path/to/speaker.wav \
  --qwen3_tts_ref_cache_dir ~/.cache/qwen3_voices

```

### Step 4: Use Cached Embeddings Directly

Skip audio processing entirely by pointing to cached files:

```bash
speech-to-speech \
  --tts qwen3 \
  --qwen3_tts_ref_spk ~/.cache/qwen3_voices/speaker.spk \
  --qwen3_tts_ref_rvq ~/.cache/qwen3_voices/speaker.rvq \
  --qwen3_tts_ref_text "Hello, this is my cloned voice."

```

## Platform-Specific Configuration

### Apple Silicon (MLX Audio)

On macOS with Apple Silicon, the handler automatically selects `mlx-audio` with built-in quantization. Backend specification is ignored:

```bash
speech-to-speech \
  --tts qwen3 \
  --qwen3_tts_backend torch \
  --qwen3_tts_ref_audio /Users/me/voice.wav

```

The `torch` backend flag has no effect on Darwin—`mlx-audio` is always used【/src/speech_to_speech/TTS/qwen3_tts_handler.py#L94-L118】.

### Linux/CUDA (GGML/Faster Backend)

Other platforms use `faster-qwen3-tts` via the `_setup_faster()` initialization path, enabling GPU-accelerated voice conversion on NVIDIA hardware.

## Real-Time Streaming Architecture

Voice conversion operates in streaming mode by default. The `streaming_chunk_size` parameter controls audio buffer sizes, enabling near-instantaneous playback as the LLM generates text and the TTS handler synthesizes speech. This architecture eliminates wait times for full utterance completion.

## Summary

- **Voice conversion** in HuggingFace speech-to-speech clones a target speaker using reference audio or cached embeddings
- **Core implementation** resides in `Qwen3TTSHandler` ([`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py)) with backend auto-detection
- **CLI arguments** in [`qwen3_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_arguments.py) expose `--qwen3_tts_ref_audio`, `--qwen3_tts_ref_spk`, and caching options
- **Pipeline wiring** occurs in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), connecting arguments to the active TTS handler
- **Three generation modes**: voice cloning, custom presets, and instruction-based voice design
- **Platform optimization**: `mlx-audio` on Apple Silicon, `faster-qwen3-tts` elsewhere

## Frequently Asked Questions

### What audio format works best for voice conversion reference files?

WAV format with 16-bit PCM encoding and 16kHz or 24kHz sample rate provides optimal results. The Qwen3-TTS handler performs internal resampling, but starting with clean, uncompressed audio minimizes artifacts. Five to thirty seconds of continuous speech captures sufficient prosodic characteristics.

### How do I speed up voice conversion on repeated runs?

Use `--qwen3_tts_ref_cache_dir` to persist speaker embeddings (`.spk`) and RVQ tokens (`.rvq`). After the first conversion, reference the cached files directly with `--qwen3_tts_ref_spk` and `--qwen3_tts_ref_rvq` instead of reprocessing the original audio. This reduces initialization from seconds to milliseconds.

### Can I convert voice without any reference audio?

Yes—use built-in speakers via `--qwen3_tts_speaker <name>` (e.g., `Aiden`, `Ada`) or supply an instruction prompt with `--qwen3_tts_instruct` for VoiceDesign models. However, true voice cloning requires either reference audio or pre-computed embeddings from a specific target speaker.

### Does voice conversion work offline?

Yes, provided the Qwen3-TTS model weights are cached locally. The pipeline performs all inference on-device; no API calls are required after initial model download. The `mlx-audio` and `faster-qwen3-tts` backends both support fully offline operation.