How to Implement Custom TTS Voice Cloning with Pocket TTS in the Speech-to-Speech Pipeline
You can clone custom voices in the Hugging Face speech-to-speech pipeline by setting the pocket_tts_voice argument to a preset name, local audio file path, or Hugging Face repository URL, which the PocketTTSHandler converts into a voice embedding via get_state_for_audio_prompt.
The huggingface/speech-to-speech repository integrates Pocket TTS (Kyutai Labs' open-source model) to enable real-time text-to-speech synthesis with custom voice cloning capabilities. By configuring the PocketTTSHandlerArguments, you can specify voice sources ranging from built-in presets to custom audio files or Hugging Face hosted embeddings. This guide covers how to implement custom TTS voice cloning with Pocket TTS using both command-line interfaces and programmatic Python configuration.
Understanding Pocket TTS Architecture
The Pocket TTS integration centers around the PocketTTSHandler class defined in src/speech_to_speech/TTS/pocket_tts_handler.py. When initializing the pipeline with module_kwargs.tts = "pocket", the factory function get_tts_handler in src/speech_to_speech/s2s_pipeline.py instantiates this handler using parameters from PocketTTSHandlerArguments located in src/speech_to_speech/arguments_classes/pocket_tts_arguments.py.
The handler manages model initialization, voice state loading, and audio streaming. It automatically resamples the native 24 kHz output to your target sample rate while processing audio in configurable blocks for real-time playback.
Voice Source Options for Custom Cloning
The pocket_tts_voice parameter accepts three distinct input types, enabling flexible voice cloning workflows:
- Preset Names — Built-in voices shipped with the model such as
"alba"or"jean"(the default). These require no additional files and load immediately. - Local Audio Files — Absolute paths to WAV files on your local filesystem, such as
"/path/to/my_voice.wav". The handler extracts voice embeddings directly from these files for instant cloning. - Hugging Face Repository Paths — References to voice embeddings stored in Hugging Face repositories using the
hf://protocol, like"hf://kyutai/tts-voices/my_voice". This enables sharing and versioning custom voices across teams.
Command-Line Voice Configuration
Configure custom voice cloning via CLI by passing the --pocket_tts_voice argument when launching the pipeline.
Use a built-in preset:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice alba
Clone from a local audio file:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice /home/user/my_voice.wav
Load a voice from a Hugging Face repository:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice "hf://kyutai/tts-voices/custom_voice"
Programmatic Implementation
For custom pipelines, instantiate PocketTTSHandlerArguments and pass it to the TTS handler factory:
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.s2s_pipeline import get_tts_handler
from threading import Event
from queue import Queue
# Configure voice cloning parameters
pocket_args = PocketTTSHandlerArguments(
pocket_tts_device="cpu", # or "cuda"/"mps"
pocket_tts_voice="hf://kyutai/tts-voices/custom_voice",
pocket_tts_sample_rate=16000,
pocket_tts_blocksize=512,
pocket_tts_max_tokens=80,
)
# Initialize pipeline queues and events
stop_event = Event()
lm_response_q = Queue()
audio_out_q = Queue()
should_listen = Event()
# Create the handler with custom voice settings
tts_handler = get_tts_handler(
module_kwargs=type("M", (), {"tts": "pocket"}),
stop_event=stop_event,
lm_response_queue=lm_response_q,
send_audio_chunks_queue=audio_out_q,
should_listen=should_listen,
chat_tts_handler_kwargs=None,
facebook_mms_tts_handler_kwargs=None,
pocket_tts_handler_kwargs=pocket_args,
kokoro_tts_handler_kwargs=None,
qwen3_tts_handler_kwargs=None,
)
Internal Voice Loading Mechanism
Inside src/speech_to_speech/TTS/pocket_tts_handler.py, the setup() method processes the voice parameter through the Pocket TTS library's abstraction layer:
# From PocketTTSHandler.setup(...)
logger.info(f"Loading voice: {voice}")
self.voice_state = self.model.get_state_for_audio_prompt(voice)
The get_state_for_audio_prompt method handles all three voice source types transparently, whether you're using a preset string, local file path, or Hugging Face URL. This unified interface simplifies voice management while supporting real-time streaming requirements.
Audio Processing and Streaming Configuration
The handler automatically manages audio format conversions and streaming buffers. Key parameters include:
- Sample Rate — Defaults to 16 kHz, with automatic resampling from the model's native 24 kHz output
- Blocksize — Configurable chunk size (default 512 samples) controlling the granularity of audio streaming to downstream components
- Device — CPU, CUDA, or MPS selection for model inference
These settings are defined in PocketTTSHandlerArguments and apply consistently across both CLI and programmatic usage.
Summary
- PocketTTSHandler in
src/speech_to_speech/TTS/pocket_tts_handler.pymanages the complete voice cloning pipeline from loading to streaming - Three voice source types are supported: preset names (
"jean","alba"), local WAV files, and Hugging Face repository paths (hf://) - Configuration occurs through
PocketTTSHandlerArgumentsinsrc/speech_to_speech/arguments_classes/pocket_tts_arguments.py - Voice embedding extraction happens via
get_state_for_audio_prompttransparently across all source types - Real-time streaming uses configurable block sizes with automatic resampling from 24 kHz to your target sample rate
Frequently Asked Questions
What audio file formats does Pocket TTS accept for voice cloning?
Pocket TTS accepts standard WAV files for local voice cloning. When providing a file path to the pocket_tts_voice argument, ensure the file is a valid WAV format that the get_state_for_audio_prompt method can process into voice embeddings.
Can I use Pocket TTS voice cloning with GPU acceleration?
Yes. Set pocket_tts_device="cuda" in the PocketTTSHandlerArguments or pass the appropriate device flag via CLI. The handler supports CPU, CUDA, and MPS backends for model inference according to the source code implementation.
How do I share custom voices across different deployments?
Store your voice embeddings in a Hugging Face repository and reference them using the hf:// protocol format (e.g., hf://kyutai/tts-voices/custom_voice). This allows consistent voice cloning across multiple environments without transferring local audio files.
What is the default voice if I don't specify a custom voice?
If no voice is specified, Pocket TTS defaults to the "jean" preset voice, as defined in the handler's argument dataclass and verified in the test suite at tests/test_cli_defaults.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →