# Speech Transformations in the Hugging Face Speech-to-Speech Pipeline

> Explore Hugging Face speech transformations including voice cloning, custom speakers, voice design, and multilingual TTS. Transform audio streams with new characteristics, languages, and styles.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: deep-dive
- Published: 2026-07-07

---

**The Hugging Face Speech-to-Speech repository supports voice cloning, custom preset speakers, voice design from text descriptions, and multilingual TTS through modular handlers that transform input audio streams into new audio with different characteristics, languages, or styles.**

The `huggingface/speech-to-speech` repository implements a modular real-time pipeline that converts input audio into transformed output audio. Built around interchangeable STT (speech-to-text), LLM, and TTS (text-to-speech) components, the system enables developers to perform voice conversion, language translation, and stylistic rewriting by routing audio through specialized handlers.

## Core Speech Transformation Types

The pipeline supports several distinct **speech transformation** modes determined by the TTS handler configuration. These transformations modify voice characteristics while preserving or altering the spoken content.

### Voice Cloning from Reference Audio

The **voice cloning** transformation reproduces the timbre of a reference speaker by extracting speaker embeddings from provided audio. This mode requires a reference audio file (`ref_audio`) and optional reference text (`ref_text`) to guide the cloning process.

In [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py), the `Qwen3TTSHandler._process_voice_clone` method handles this transformation. The handler processes the reference audio to generate consistent speaker embeddings that synthesize speech matching the original voice characteristics.

```python
args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"

```

### Custom Voice Selection from Presets

The **custom-voice** transformation generates speech using predefined speakers from a catalog (e.g., "Aiden", "Maya"). This mode selects speaker embeddings from the model's internal library rather than extracting them from reference audio.

The `Qwen3TTSHandler._process_custom_voice` method in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) implements this through the `_resolve_speaker` helper function. Users specify the speaker via the `qwen3_tts_speaker` argument.

```python
args.qwen3_tts_handler_kwargs.speaker = "Aiden"
args.module_kwargs.tts = "qwen3"

```

### Voice Design via Text Instructions

The **voice design** transformation synthesizes novel voices from textual descriptions without requiring reference audio. Users provide an instruction prompt (`instruct`) describing desired characteristics such as "a calm, deep voice with slight echo."

The `Qwen3TTSHandler._process_voice_design` method processes these instructions to create new voice embeddings on-the-fly. This method is located at line 851 of [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py).

```python
args.qwen3_tts_handler_kwargs.instruct = "a calm, deep voice with slight echo"
args.module_kwargs.tts = "qwen3"

```

### Generic and Specialized TTS Handlers

Beyond the Qwen3-TTS handler, the repository provides several specialized handlers for specific use cases:

- **ChatTTS**: Provides a default generic voice using randomly sampled speaker embeddings. Implemented in [`src/speech_to_speech/TTS/chatTTS_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/chatTTS_handler.py) via the `ChatTTSHandler.setup` method.
- **Pocket TTS**: Lightweight multilingual TTS available through the `speech-to-speech[TTS,pocket]` optional dependency. Located in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py).
- **Kokoro TTS**: Japanese-focused TTS with voice-style control via the `kokoro_tts_speaker` parameter. Located in [`src/speech_to_speech/TTS/kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/kokoro_handler.py).
- **Facebook MMS TTS**: Multilingual streaming TTS from Meta, available via `speech-to-speech[facebook-mms]`. Located in [`src/speech_to_speech/TTS/facebookmms_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/facebookmms_handler.py).

## Pipeline Architecture for Advanced Transformations

The modular architecture enables complex **speech transformations** by combining TTS handlers with STT and LLM backends.

### STT and LLM Integration

The pipeline supports multiple STT backends including Whisper, Faster-Whisper, Paraformer, MLX-Audio-Whisper, and Parakeet-TDT. These feed into LLM backends (Transformers, MLX-LM, OpenAI-compatible APIs) that can perform:

- **Voice conversion**: Preserve content while changing the speaker identity
- **Language translation**: STT → LLM translation → TTS in target language
- **Stylistic rewriting**: LLM modifies the transcript (e.g., "make it formal") before synthesis
- **Realtime voice-over**: Live microphone input processed through VAD → STT → LLM → TTS with speculative turn handling for low latency

All configurations are managed through argument classes in `src/speech_to_speech/arguments_classes/`, which define parameters for each handler type.

## Configuration and Usage Examples

### Complete Voice Cloning Setup

The following example demonstrates configuring the pipeline for voice cloning on Apple Silicon:

```python
from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args, build_pipeline
from speech_to_speech.utils.thread_manager import ThreadManager

# Parse CLI args or construct them manually

args = parse_arguments()

# Configure for voice cloning

args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"

# Prepare arguments for all handlers

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Build and start pipeline

pipeline_manager: ThreadManager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    initialize_queues_and_events(),
)

pipeline_manager.start()

```

### Language Translation Pipeline

To combine **speech transformation** with translation:

```python

# LLM configuration for translation

args.language_model_handler_kwargs.model_name = "facebook/opt-1.3b"
args.language_model_handler_kwargs.chat_prompt = "Translate English to French and respond."

# TTS configuration for French output

args.qwen3_tts_handler_kwargs.language = "fr"
args.module_kwargs.tts = "qwen3"

```

### Quick-Start Generic TTS

For immediate deployment without reference audio:

```python
args.module_kwargs.tts = "chatTTS"

```

This instantiates `ChatTTSHandler`, which automatically selects a random speaker embedding from the generic multilingual model.

## Summary

- **Voice cloning** reproduces speaker timbre from reference audio using `Qwen3TTSHandler._process_voice_clone` in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py).
- **Custom voices** select from preset speakers like "Aiden" via the `qwen3_tts_speaker` parameter and `Qwen3TTSHandler._process_custom_voice`.
- **Voice design** creates novel voices from text descriptions using `Qwen3TTSHandler._process_voice_design` and the `instruct` parameter.
- **Specialized handlers** include ChatTTS, Pocket TTS, Kokoro TTS, and Facebook MMS TTS for specific languages and deployment constraints.
- **Advanced transformations** combine STT, LLM, and TTS components to enable real-time translation, voice conversion, and stylistic rewriting.
- All handlers expose a common `process()` method returning `int16` PCM chunks, enabling seamless integration with output streamers like `LocalAudioStreamer` or `WebSocketStreamer`.

## Frequently Asked Questions

### What audio format is required for voice cloning?

Voice cloning requires a **reference audio file** (`ref_audio`) in a standard format (typically WAV) and optionally **reference text** (`ref_text`) containing the transcript of the reference audio. The `Qwen3TTSHandler` extracts speaker embeddings from this audio to synthesize matching voice characteristics.

### Can I use the speech-to-speech pipeline for real-time voice conversion?

Yes, the pipeline supports **real-time voice-over** by processing live microphone input through Voice Activity Detection (VAD) → STT → LLM → TTS with speculative turn handling for low latency. All TTS handlers return audio as streams of `int16` PCM chunks suitable for real-time playback.

### How do I switch between different TTS backends?

Set `args.module_kwargs.tts` to the desired handler name (`"qwen3"`, `"chatTTS"`, `"pocket"`, `"kokoro"`, or `"facebook_mms"`) and ensure the corresponding handler arguments are prepared via `prepare_all_args()`. The pipeline automatically instantiates the correct handler class from `src/speech_to_speech/TTS/` based on this configuration.

### Does the pipeline support multilingual speech transformations?

Yes, the **Qwen3-TTS** handler supports multiple languages through the `language` parameter, while specialized handlers like **Kokoro TTS** (Japanese) and **Facebook MMS TTS** provide additional multilingual capabilities. When combined with LLM translation, the pipeline can perform cross-lingual voice conversion (e.g., English input to French output with cloned voice characteristics).