# How to Perform Voice Conversion with Speech-to-Speech: A Complete Pipeline Guide

> Master voice conversion with speech-to-speech. Learn the four-stage pipeline VAD STT LLM TTS for real-time voice transformation and synthesize any target voice.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: tutorial
- Published: 2026-08-01

---

**Voice conversion with speech-to-speech is achieved through a four-stage pipeline that processes audio through Voice Activity Detection (VAD), Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS), allowing real-time transformation of any input voice into a synthesized target voice.**

The huggingface/speech-to-speech repository provides an open-source framework for real-time voice conversion. This system implements a modular pipeline that captures spoken audio, transcribes it, generates a textual response, and synthesizes it into a completely different voice. By leveraging configurable backends for each stage, you can perform voice conversion with speech-to-speech using either command-line interfaces or direct Python APIs.

## Understanding the Four-Stage Voice Conversion Architecture

The pipeline operates as a series of independent threads connected by thread-safe queues, as orchestrated in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py). This architecture minimizes latency while processing audio streams end-to-end.

### Stage 1: Voice Activity Detection with Silero VAD

The process begins with **Voice Activity Detection (VAD)** using Silero VAD v5, implemented in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py). This component detects speech boundaries and turn-taking events, ensuring only relevant audio segments proceed to transcription.

### Stage 2: Speech-to-Text Transcription

Next, the **Speech-to-Text (STT)** stage converts audio to text. The default backend is Parakeet TDT, though the system supports alternatives including Whisper, Faster-Whisper, Lightning-Whisper-MLX, MLX-Audio-Whisper, and Paraformer. These implementations reside in the `src/speech_to_speech/STT/` directory and are selected via the `S2SPipeline` configuration.

### Stage 3: Language Model Processing

The transcribed text feeds into a **Language Model (LLM)** for response generation. According to [`src/speech_to_speech/pipeline/handler_types.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/handler_types.py), supported backends include OpenAI-compatible APIs, Hugging Face Transformers models, and mlx-lm for Apple Silicon devices.

### Stage 4: Text-to-Speech Voice Synthesis

Finally, **Text-to-Speech (TTS)** synthesizes the output audio. The default Qwen3-TTS backend in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) can be swapped with alternatives like Kokoro-82M, Pocket TTS, ChatTTS, or Facebook MMS. By selecting different TTS models or speaker checkpoints, you control the final converted voice characteristics.

## Running the Real-Time Voice Conversion Server

To start voice conversion immediately, install the package and launch the built-in server.

```bash
pip install speech-to-speech

export OPENAI_API_KEY=...
speech-to-speech

```

The server initializes the complete VAD → STT → LLM → TTS pipeline and exposes a WebSocket endpoint at `ws://localhost:8765/v1/realtime`. This endpoint accepts audio streams and returns synthesized voice output in real-time.

## Testing Voice Conversion with a Client

Connect a microphone and speaker to test the conversion loop using the provided demo script.

```bash
python scripts/listen_and_play_realtime.py \
    --host 127.0.0.1 --port 8765

```

This script records from your microphone, streams audio to the server, and plays back the converted voice, demonstrating the complete voice conversion workflow.

## Implementing Voice Conversion Programmatically

For custom integrations, instantiate the `S2SPipeline` class directly to configure specific backends and model checkpoints.

```python
from speech_to_speech.s2s_pipeline import S2SPipeline

pipeline = S2SPipeline(
    vad="silero",
    stt="parakeet-tdt",
    llm_backend="responses-api",
    tts="pocket",
    tts_model_name="kyutai/pocket-tts-base",
    tts_speaker="female",
)

output_wav = pipeline.run("input.wav")

```

This approach allows you to mix and match components—such as retaining the default STT while switching to Pocket TTS for a distinct voice output—enabling flexible voice conversion pipelines.

## Customizing Voices with Model Checkpoints

To convert speech to a specific target voice using Qwen3-TTS, specify a custom checkpoint and speaker identity via CLI flags.

```bash
speech-to-speech \
    --tts qwen3 \
    --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
    --qwen3_tts_speaker "Aiden"

```

The `--qwen3_tts_speaker` parameter selects the target voice persona, while `--qwen3_tts_model_name` accepts any fine-tuned checkpoint for personalized voice conversion.

## Summary

- Voice conversion with speech-to-speech utilizes a four-stage pipeline: VAD, STT, LLM, and TTS, connected via thread-safe queues for low-latency processing.
- The default implementation uses Silero VAD v5, Parakeet TDT for transcription, and Qwen3-TTS for synthesis, as defined in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py).
- Launch the real-time server with the `speech-to-speech` CLI command, which exposes a WebSocket endpoint at `ws://localhost:8765/v1/realtime`.
- Customize voice outputs by swapping TTS backends (Kokoro-82M, Pocket TTS, ChatTTS) or loading specific speaker checkpoints via the `S2SPipeline` API or CLI flags.
- Test the complete conversion loop using [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) to stream microphone input and receive synthesized audio responses.

## Frequently Asked Questions

### What is the default TTS backend for voice conversion in speech-to-speech?

The default TTS backend is **Qwen3-TTS**, implemented in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py). This model synthesizes natural-sounding speech and supports speaker selection via the `--qwen3_tts_speaker` CLI flag. You can replace it with alternatives like Kokoro-82M or Pocket TTS by specifying the `--tts` parameter.

### Can I use speech-to-speech for voice conversion without an internet connection?

Yes, the pipeline supports fully local execution by selecting offline-capable backends. Use `mlx-lm` for the LLM stage on Apple Silicon, and local Whisper or Parakeet checkpoints for STT. For TTS, models like Kokoro-82M or fine-tuned Qwen3-TTS checkpoints can run entirely on local hardware without API calls.

### How does the pipeline handle real-time latency during voice conversion?

The architecture employs **thread-safe queues** to connect each stage (VAD, STT, LLM, TTS) in separate threads. This concurrent processing allows the system to transcribe incoming audio while simultaneously synthesizing previous segments, minimizing end-to-end latency for real-time voice conversion applications.

### Where can I find the voice activity detection implementation?

The VAD logic resides in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py) and [`src/speech_to_speech/VAD/vad_iterator.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_iterator.py). These files implement the Silero VAD v5 model, which detects speech boundaries and manages turn-taking by yielding segmented audio chunks to the STT handler.