How to Configure Custom Voice Cloning with Pocket TTS in the Speech-to-Speech Pipeline
You can configure custom voice cloning with Pocket TTS by setting the pocket_tts_voice argument to a preset name, local audio file path, or Hugging Face repository URL, which the PocketTTSHandler processes through the get_state_for_audio_prompt method.
The huggingface/speech-to-speech repository integrates Kyutai Labs' Pocket TTS model through a dedicated handler that supports zero-shot voice cloning. When you select Pocket TTS as your text-to-speech backend, the pipeline instantiates PocketTTSHandler from src/speech_to_speech/TTS/pocket_tts_handler.py using configuration parameters defined in PocketTTSHandlerArguments.
Understanding Pocket TTS Integration
The Pocket TTS integration follows the modular architecture of the speech-to-speech pipeline. When module_kwargs.tts is set to "pocket", the factory function get_tts_handler in src/speech_to_speech/s2s_pipeline.py initializes the handler with your custom arguments.
The handler manages the entire lifecycle of voice synthesis: loading the Kyutai model onto your specified device, preparing the voice embedding, generating audio at 24 kHz, and resampling to your target sample rate for real-time streaming.
Voice Configuration Options
The voice parameter in PocketTTSHandlerArguments accepts three distinct input types, giving you flexibility in how you source your cloning target.
Preset Voices
Pocket TTS ships with built-in voice presets including "alba" and "jean". The default voice is "jean" as specified in the handler's docstring and argument defaults. These require no external files and load immediately from the model weights.
Local Audio Files
You can clone any speaker by providing a path to a local WAV file. The handler extracts a voice embedding from this audio and uses it to condition the generation. The source audio should be clear speech for optimal cloning quality.
Hugging Face Repository Paths
For shared or version-controlled voices, pass a Hugging Face repository path using the hf:// protocol. The handler downloads the voice embedding from the specified repository location, enabling team collaboration and reproducible deployments.
Configuring Voice Cloning via Command Line
The speech-to-speech CLI exposes the pocket_tts_voice argument to configure your desired voice source. Here are the three supported patterns:
Use a built-in preset voice:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice alba
Clone a custom voice from a local audio file:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice /home/user/my_voice.wav
Load a voice from a Hugging Face repository:
python -m speech_to_speech.main --tts pocket \
--pocket_tts_voice "hf://kyutai/tts-voices/custom_voice"
Configuring Voice Cloning Programmatically
For custom pipelines or research applications, import PocketTTSHandlerArguments from src/speech_to_speech/arguments_classes/pocket_tts_arguments.py and instantiate the handler directly.
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.s2s_pipeline import get_tts_handler
from queue import Queue
from threading import Event
# Configure the voice and processing parameters
pocket_args = PocketTTSHandlerArguments(
pocket_tts_device="cuda", # or "cpu", "mps"
pocket_tts_voice="hf://kyutai/tts-voices/custom_voice",
pocket_tts_sample_rate=16000,
pocket_tts_blocksize=512,
pocket_tts_max_tokens=80,
)
# Create required pipeline components
stop_event = Event()
lm_response_q = Queue()
audio_out_q = Queue()
should_listen = Event()
# Instantiate the handler
tts_handler = get_tts_handler(
module_kwargs=type("M", (), {"tts": "pocket"}),
stop_event=stop_event,
lm_response_queue=lm_response_q,
send_audio_chunks_queue=audio_out_q,
should_listen=should_listen,
chat_tts_handler_kwargs=None,
facebook_mms_tts_handler_kwargs=None,
pocket_tts_handler_kwargs=pocket_args,
kokoro_tts_handler_kwargs=None,
qwen3_tts_handler_kwargs=None,
)
Technical Implementation Details
The voice cloning mechanism centers on the setup() method in src/speech_to_speech/TTS/pocket_tts_handler.py. When the handler initializes, it processes the voice argument through the following code:
# Inside PocketTTSHandler.setup(...)
logger.info(f"Loading voice: {voice}")
self.voice_state = self.model.get_state_for_audio_prompt(voice)
The get_state_for_audio_prompt method (provided by the underlying pocket_tts library) abstracts the loading logic for all three voice source types. It returns a voice state tensor that conditions the autoregressive generation throughout the streaming process.
Audio Processing Pipeline:
- Native generation: The model generates audio at 24 kHz
- Resampling: The handler automatically resamples to your configured
sample_rate(default 16 kHz) - Streaming: Audio emits in blocks of
blocksizesamples (default 512) to minimize latency for real-time playback
Summary
- Custom voice cloning with Pocket TTS requires setting the
pocket_tts_voiceargument to either a preset name, local WAV path, or Hugging Face repository URL - The
PocketTTSHandlerclass insrc/speech_to_speech/TTS/pocket_tts_handler.pymanages voice loading throughget_state_for_audio_promptand handles resampling from 24 kHz to your target sample rate - Configuration is controlled via
PocketTTSHandlerArgumentsinsrc/speech_to_speech/arguments_classes/pocket_tts_arguments.py, exposing parameters for device, sample rate, blocksize, and max tokens - The default voice is
"jean"when no custom voice is specified - The handler streams audio in 512-sample blocks by default, enabling real-time speech-to-speech conversation
Frequently Asked Questions
What voice formats does Pocket TTS support for cloning?
Pocket TTS accepts three voice identifier formats: built-in string presets like "alba" or "jean", absolute file paths to local WAV files (e.g., /path/to/voice.wav), and Hugging Face repository URLs using the hf:// protocol (e.g., hf://kyutai/tts-voices/custom_voice). The handler automatically detects the format and processes it through get_state_for_audio_prompt.
How do I change the audio sample rate in Pocket TTS?
Set the pocket_tts_sample_rate argument when configuring PocketTTSHandlerArguments. While the model natively generates at 24 kHz, the handler resamples to your specified rate (default 16 kHz) before queuing audio for playback. This parameter is available in both CLI (--pocket_tts_sample_rate) and programmatic interfaces.
Can I use Pocket TTS with GPU acceleration?
Yes. Set pocket_tts_device="cuda" in your PocketTTSHandlerArguments or pass --pocket_tts_device cuda via CLI. The handler calls .to(device) on the model during initialization in src/speech_to_speech/TTS/pocket_tts_handler.py, supporting CUDA, MPS (Apple Silicon), and CPU backends.
Where is the default voice configured in the source code?
The default voice "jean" is defined in src/speech_to_speech/arguments_classes/pocket_tts_arguments.py within the PocketTTSHandlerArguments dataclass. The test suite in tests/test_cli_defaults.py verifies that this default is correctly wired into the pipeline when no custom voice is specified.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →