How to Perform Voice Conversion with Speech-to-Speech: A Complete Pipeline Guide

Voice conversion with speech-to-speech is achieved through a four-stage pipeline that processes audio through Voice Activity Detection (VAD), Speech-to-Text (STT), a Language Model (LLM), and Text-to-Speech (TTS), allowing real-time transformation of any input voice into a synthesized target voice.

The huggingface/speech-to-speech repository provides an open-source framework for real-time voice conversion. This system implements a modular pipeline that captures spoken audio, transcribes it, generates a textual response, and synthesizes it into a completely different voice. By leveraging configurable backends for each stage, you can perform voice conversion with speech-to-speech using either command-line interfaces or direct Python APIs.

Understanding the Four-Stage Voice Conversion Architecture

The pipeline operates as a series of independent threads connected by thread-safe queues, as orchestrated in src/speech_to_speech/s2s_pipeline.py. This architecture minimizes latency while processing audio streams end-to-end.

Stage 1: Voice Activity Detection with Silero VAD

The process begins with Voice Activity Detection (VAD) using Silero VAD v5, implemented in src/speech_to_speech/VAD/vad_handler.py. This component detects speech boundaries and turn-taking events, ensuring only relevant audio segments proceed to transcription.

Stage 2: Speech-to-Text Transcription

Next, the Speech-to-Text (STT) stage converts audio to text. The default backend is Parakeet TDT, though the system supports alternatives including Whisper, Faster-Whisper, Lightning-Whisper-MLX, MLX-Audio-Whisper, and Paraformer. These implementations reside in the src/speech_to_speech/STT/ directory and are selected via the S2SPipeline configuration.

Stage 3: Language Model Processing

The transcribed text feeds into a Language Model (LLM) for response generation. According to src/speech_to_speech/pipeline/handler_types.py, supported backends include OpenAI-compatible APIs, Hugging Face Transformers models, and mlx-lm for Apple Silicon devices.

Stage 4: Text-to-Speech Voice Synthesis

Finally, Text-to-Speech (TTS) synthesizes the output audio. The default Qwen3-TTS backend in src/speech_to_speech/TTS/qwen3_tts_handler.py can be swapped with alternatives like Kokoro-82M, Pocket TTS, ChatTTS, or Facebook MMS. By selecting different TTS models or speaker checkpoints, you control the final converted voice characteristics.

Running the Real-Time Voice Conversion Server

To start voice conversion immediately, install the package and launch the built-in server.

pip install speech-to-speech

export OPENAI_API_KEY=...
speech-to-speech

The server initializes the complete VAD → STT → LLM → TTS pipeline and exposes a WebSocket endpoint at ws://localhost:8765/v1/realtime. This endpoint accepts audio streams and returns synthesized voice output in real-time.

Testing Voice Conversion with a Client

Connect a microphone and speaker to test the conversion loop using the provided demo script.

python scripts/listen_and_play_realtime.py \
    --host 127.0.0.1 --port 8765

This script records from your microphone, streams audio to the server, and plays back the converted voice, demonstrating the complete voice conversion workflow.

Implementing Voice Conversion Programmatically

For custom integrations, instantiate the S2SPipeline class directly to configure specific backends and model checkpoints.

from speech_to_speech.s2s_pipeline import S2SPipeline

pipeline = S2SPipeline(
    vad="silero",
    stt="parakeet-tdt",
    llm_backend="responses-api",
    tts="pocket",
    tts_model_name="kyutai/pocket-tts-base",
    tts_speaker="female",
)

output_wav = pipeline.run("input.wav")

This approach allows you to mix and match components—such as retaining the default STT while switching to Pocket TTS for a distinct voice output—enabling flexible voice conversion pipelines.

Customizing Voices with Model Checkpoints

To convert speech to a specific target voice using Qwen3-TTS, specify a custom checkpoint and speaker identity via CLI flags.

speech-to-speech \
    --tts qwen3 \
    --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
    --qwen3_tts_speaker "Aiden"

The --qwen3_tts_speaker parameter selects the target voice persona, while --qwen3_tts_model_name accepts any fine-tuned checkpoint for personalized voice conversion.

Summary

  • Voice conversion with speech-to-speech utilizes a four-stage pipeline: VAD, STT, LLM, and TTS, connected via thread-safe queues for low-latency processing.
  • The default implementation uses Silero VAD v5, Parakeet TDT for transcription, and Qwen3-TTS for synthesis, as defined in src/speech_to_speech/s2s_pipeline.py.
  • Launch the real-time server with the speech-to-speech CLI command, which exposes a WebSocket endpoint at ws://localhost:8765/v1/realtime.
  • Customize voice outputs by swapping TTS backends (Kokoro-82M, Pocket TTS, ChatTTS) or loading specific speaker checkpoints via the S2SPipeline API or CLI flags.
  • Test the complete conversion loop using scripts/listen_and_play_realtime.py to stream microphone input and receive synthesized audio responses.

Frequently Asked Questions

What is the default TTS backend for voice conversion in speech-to-speech?

The default TTS backend is Qwen3-TTS, implemented in src/speech_to_speech/TTS/qwen3_tts_handler.py. This model synthesizes natural-sounding speech and supports speaker selection via the --qwen3_tts_speaker CLI flag. You can replace it with alternatives like Kokoro-82M or Pocket TTS by specifying the --tts parameter.

Can I use speech-to-speech for voice conversion without an internet connection?

Yes, the pipeline supports fully local execution by selecting offline-capable backends. Use mlx-lm for the LLM stage on Apple Silicon, and local Whisper or Parakeet checkpoints for STT. For TTS, models like Kokoro-82M or fine-tuned Qwen3-TTS checkpoints can run entirely on local hardware without API calls.

How does the pipeline handle real-time latency during voice conversion?

The architecture employs thread-safe queues to connect each stage (VAD, STT, LLM, TTS) in separate threads. This concurrent processing allows the system to transcribe incoming audio while simultaneously synthesizing previous segments, minimizing end-to-end latency for real-time voice conversion applications.

Where can I find the voice activity detection implementation?

The VAD logic resides in src/speech_to_speech/VAD/vad_handler.py and src/speech_to_speech/VAD/vad_iterator.py. These files implement the Silero VAD v5 model, which detects speech boundaries and manages turn-taking by yielding segmented audio chunks to the STT handler.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →