Speech Transformations in the Hugging Face Speech-to-Speech Pipeline

The Hugging Face Speech-to-Speech repository supports voice cloning, custom preset speakers, voice design from text descriptions, and multilingual TTS through modular handlers that transform input audio streams into new audio with different characteristics, languages, or styles.

The huggingface/speech-to-speech repository implements a modular real-time pipeline that converts input audio into transformed output audio. Built around interchangeable STT (speech-to-text), LLM, and TTS (text-to-speech) components, the system enables developers to perform voice conversion, language translation, and stylistic rewriting by routing audio through specialized handlers.

Core Speech Transformation Types

The pipeline supports several distinct speech transformation modes determined by the TTS handler configuration. These transformations modify voice characteristics while preserving or altering the spoken content.

Voice Cloning from Reference Audio

The voice cloning transformation reproduces the timbre of a reference speaker by extracting speaker embeddings from provided audio. This mode requires a reference audio file (ref_audio) and optional reference text (ref_text) to guide the cloning process.

In src/speech_to_speech/TTS/qwen3_tts_handler.py, the Qwen3TTSHandler._process_voice_clone method handles this transformation. The handler processes the reference audio to generate consistent speaker embeddings that synthesize speech matching the original voice characteristics.

args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"

Custom Voice Selection from Presets

The custom-voice transformation generates speech using predefined speakers from a catalog (e.g., "Aiden", "Maya"). This mode selects speaker embeddings from the model's internal library rather than extracting them from reference audio.

The Qwen3TTSHandler._process_custom_voice method in src/speech_to_speech/TTS/qwen3_tts_handler.py implements this through the _resolve_speaker helper function. Users specify the speaker via the qwen3_tts_speaker argument.

args.qwen3_tts_handler_kwargs.speaker = "Aiden"
args.module_kwargs.tts = "qwen3"

Voice Design via Text Instructions

The voice design transformation synthesizes novel voices from textual descriptions without requiring reference audio. Users provide an instruction prompt (instruct) describing desired characteristics such as "a calm, deep voice with slight echo."

The Qwen3TTSHandler._process_voice_design method processes these instructions to create new voice embeddings on-the-fly. This method is located at line 851 of src/speech_to_speech/TTS/qwen3_tts_handler.py.

args.qwen3_tts_handler_kwargs.instruct = "a calm, deep voice with slight echo"
args.module_kwargs.tts = "qwen3"

Generic and Specialized TTS Handlers

Beyond the Qwen3-TTS handler, the repository provides several specialized handlers for specific use cases:

Pipeline Architecture for Advanced Transformations

The modular architecture enables complex speech transformations by combining TTS handlers with STT and LLM backends.

STT and LLM Integration

The pipeline supports multiple STT backends including Whisper, Faster-Whisper, Paraformer, MLX-Audio-Whisper, and Parakeet-TDT. These feed into LLM backends (Transformers, MLX-LM, OpenAI-compatible APIs) that can perform:

  • Voice conversion: Preserve content while changing the speaker identity
  • Language translation: STT → LLM translation → TTS in target language
  • Stylistic rewriting: LLM modifies the transcript (e.g., "make it formal") before synthesis
  • Realtime voice-over: Live microphone input processed through VAD → STT → LLM → TTS with speculative turn handling for low latency

All configurations are managed through argument classes in src/speech_to_speech/arguments_classes/, which define parameters for each handler type.

Configuration and Usage Examples

Complete Voice Cloning Setup

The following example demonstrates configuring the pipeline for voice cloning on Apple Silicon:

from speech_to_speech.s2s_pipeline import parse_arguments, prepare_all_args, build_pipeline
from speech_to_speech.utils.thread_manager import ThreadManager

# Parse CLI args or construct them manually

args = parse_arguments()

# Configure for voice cloning

args.qwen3_tts_handler_kwargs.ref_audio = "samples/my_reference.wav"
args.qwen3_tts_handler_kwargs.ref_text = "I love coffee."
args.qwen3_tts_handler_kwargs.language = "en"
args.module_kwargs.tts = "qwen3"

# Prepare arguments for all handlers

prepare_all_args(
    args.module_kwargs,
    args.whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
)

# Build and start pipeline

pipeline_manager: ThreadManager = build_pipeline(
    args.module_kwargs,
    args.socket_receiver_kwargs,
    args.socket_sender_kwargs,
    args.websocket_streamer_kwargs,
    args.vad_handler_kwargs,
    args.whisper_stt_handler_kwargs,
    args.faster_whisper_stt_handler_kwargs,
    args.paraformer_stt_handler_kwargs,
    args.mlx_audio_whisper_stt_handler_kwargs,
    args.parakeet_tdt_stt_handler_kwargs,
    args.language_model_handler_kwargs,
    args.responses_api_language_model_handler_kwargs,
    args.chat_tts_handler_kwargs,
    args.facebook_mms_tts_handler_kwargs,
    args.pocket_tts_handler_kwargs,
    args.kokoro_tts_handler_kwargs,
    args.qwen3_tts_handler_kwargs,
    initialize_queues_and_events(),
)

pipeline_manager.start()

Language Translation Pipeline

To combine speech transformation with translation:


# LLM configuration for translation

args.language_model_handler_kwargs.model_name = "facebook/opt-1.3b"
args.language_model_handler_kwargs.chat_prompt = "Translate English to French and respond."

# TTS configuration for French output

args.qwen3_tts_handler_kwargs.language = "fr"
args.module_kwargs.tts = "qwen3"

Quick-Start Generic TTS

For immediate deployment without reference audio:

args.module_kwargs.tts = "chatTTS"

This instantiates ChatTTSHandler, which automatically selects a random speaker embedding from the generic multilingual model.

Summary

  • Voice cloning reproduces speaker timbre from reference audio using Qwen3TTSHandler._process_voice_clone in src/speech_to_speech/TTS/qwen3_tts_handler.py.
  • Custom voices select from preset speakers like "Aiden" via the qwen3_tts_speaker parameter and Qwen3TTSHandler._process_custom_voice.
  • Voice design creates novel voices from text descriptions using Qwen3TTSHandler._process_voice_design and the instruct parameter.
  • Specialized handlers include ChatTTS, Pocket TTS, Kokoro TTS, and Facebook MMS TTS for specific languages and deployment constraints.
  • Advanced transformations combine STT, LLM, and TTS components to enable real-time translation, voice conversion, and stylistic rewriting.
  • All handlers expose a common process() method returning int16 PCM chunks, enabling seamless integration with output streamers like LocalAudioStreamer or WebSocketStreamer.

Frequently Asked Questions

What audio format is required for voice cloning?

Voice cloning requires a reference audio file (ref_audio) in a standard format (typically WAV) and optionally reference text (ref_text) containing the transcript of the reference audio. The Qwen3TTSHandler extracts speaker embeddings from this audio to synthesize matching voice characteristics.

Can I use the speech-to-speech pipeline for real-time voice conversion?

Yes, the pipeline supports real-time voice-over by processing live microphone input through Voice Activity Detection (VAD) → STT → LLM → TTS with speculative turn handling for low latency. All TTS handlers return audio as streams of int16 PCM chunks suitable for real-time playback.

How do I switch between different TTS backends?

Set args.module_kwargs.tts to the desired handler name ("qwen3", "chatTTS", "pocket", "kokoro", or "facebook_mms") and ensure the corresponding handler arguments are prepared via prepare_all_args(). The pipeline automatically instantiates the correct handler class from src/speech_to_speech/TTS/ based on this configuration.

Does the pipeline support multilingual speech transformations?

Yes, the Qwen3-TTS handler supports multiple languages through the language parameter, while specialized handlers like Kokoro TTS (Japanese) and Facebook MMS TTS provide additional multilingual capabilities. When combined with LLM translation, the pipeline can perform cross-lingual voice conversion (e.g., English input to French output with cloned voice characteristics).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →