How to Use Custom Voice Cloning with Pocket TTS in Hugging Face Speech-to-Speech
Pocket TTS supports real-time voice cloning from local audio files or Hugging Face repositories by passing a path or URL to --pocket_tts_voice, which extracts a voice state via get_state_for_audio_prompt() in the PocketTTSHandler class.
The huggingface/speech-to-speech repository provides a modular pipeline for real-time speech-to-speech translation, featuring a Pocket TTS backend capable of cloning arbitrary voices from audio references. This voice cloning functionality is implemented in the PocketTTSHandler class, which extracts voice characteristics from presets, local files, or remote repositories using the get_state_for_audio_prompt() method. Understanding how to configure the pocket_tts_voice parameter enables you to generate personalized speech output tailored to specific speakers.
Understanding the Pocket TTS Architecture
The PocketTTSHandler Class
Located in src/speech_to_speech/TTS/pocket_tts_handler.py, the PocketTTSHandler inherits from BaseHandler[TTSIn, TTSOut] and manages the voice cloning lifecycle. During initialization, the handler loads the Pocket TTS model onto the specified device and prepares the voice state that will be used for all subsequent text-to-speech generation.
Voice State Extraction
The core cloning mechanism occurs in get_state_for_audio_prompt(), which processes the input specified by pocket_tts_voice. This method accepts preset names (such as alba, marius, javert, jean, fantine, cosette, eponine, or azelma), absolute paths to local WAV files, or Hugging Face repository URLs in the format hf://kyutai/tts-voices/.... The extracted voice state is cached for efficient streaming during the process() method, which yields int16 audio blocks resampled from the model's native 24 kHz to the target sample rate using scipy.signal.resample_poly.
Command-Line Voice Cloning Configuration
Cloning from a Hugging Face Repository
To clone a voice hosted on the Hugging Face Hub, pass the repository URL prefixed with hf:// to the --pocket_tts_voice argument. The handler automatically downloads the reference audio and metadata, extracting the voice state before generating speech.
python s2s_pipeline.py \
--tts pocket \
--pocket_tts_voice hf://kyutai/tts-voices/my_custom_voice \
--pocket_tts_device cuda \
--pocket_tts_sample_rate 16000
Cloning from a Local Audio File
For private voice samples stored locally, provide the absolute path to the reference WAV file. The handler reads the waveform directly and computes the voice state without network requests.
python s2s_pipeline.py \
--tts pocket \
--pocket_tts_voice /absolute/path/to/my_reference.wav \
--pocket_tts_device cpu
Programmatic Voice Cloning in Python
For integration into custom applications, instantiate PocketTTSHandler directly with your voice source. The handler requires a should_listen event for pipeline coordination and accepts the same voice parameters as the CLI.
from threading import Event
from speech_to_speech.TTS.pocket_tts_handler import PocketTTSHandler
from speech_to_speech.pipeline.handler_types import TTSIn
# Initialize the handler with custom voice
should_listen = Event()
handler = PocketTTSHandler(
should_listen,
device="cuda",
voice="hf://kyutai/tts-voices/custom_voice",
sample_rate=16000,
blocksize=512,
)
# Setup loads the model and extracts voice state
handler.setup(
should_listen,
device="cuda",
voice="hf://kyutai/tts-voices/custom_voice",
)
# Generate speech from text
tts_input = TTSIn(text="This is a cloned voice speaking.", language_code="en")
for audio_block in handler.process(tts_input):
# audio_block is int16 numpy array of length blocksize
play_audio(audio_block) # Replace with your audio output logic
Key Configuration Parameters
The following arguments control voice cloning behavior when using the Pocket TTS backend:
pocket_tts_voice: Specifies the voice source. Accepts preset names (jean,alba, etc.), local file paths, orhf://URLs. Default:jean.pocket_tts_device: Compute device for model inference. Options:cpu,cuda,mps. Default:cpu.pocket_tts_sample_rate: Output audio sample rate. The model generates at 24 kHz internally, then resamples to this value. Default:16000.
Summary
- Pocket TTS voice cloning is implemented in
src/speech_to_speech/TTS/pocket_tts_handler.pyvia thePocketTTSHandlerclass. - Clone voices by setting
--pocket_tts_voiceto a preset name, local file path, orhf://repository URL. - The
get_state_for_audio_prompt()method extracts voice characteristics from the reference audio during handler setup. - Audio output is automatically resampled from 24 kHz to your specified
pocket_tts_sample_rateusingscipy.signal.resample_poly. - The handler integrates with the speculative turn tracker and cancel scope to ensure only fresh voice-cloned audio is generated.
Frequently Asked Questions
What audio file formats are supported for custom voice cloning?
The PocketTTSHandler reads reference audio using standard audio I/O libraries available in the environment. While the documentation emphasizes WAV files, any format supported by the underlying audio processing dependencies can be loaded from local paths or Hugging Face repositories. The reference audio should contain clear speech samples for optimal voice extraction.
Can I switch between different cloned voices at runtime without restarting the pipeline?
Yes, you can change the voice by passing a different --pocket_tts_voice value when restarting the pipeline, or programmatically by creating a new handler instance with the desired voice parameter. The setup() method re-initializes the voice state via get_state_for_audio_prompt() for the new source, allowing dynamic switching between presets, local files, or remote repositories.
How does the handler manage audio resampling for voice-cloned output?
Pocket TTS generates audio at a fixed 24 kHz sample rate internally. The handler calculates a resampling ratio based on your configured pocket_tts_sample_rate and applies scipy.signal.resample_poly to produce audio blocks matching the pipeline's default blocksize of 512 samples. This ensures compatible output regardless of the target sample rate.
Does the Pocket TTS handler support concurrent voice cloning requests?
The handler inherits cancellation logic from BaseHandler that interacts with a SpeculativeTurnTracker and CancelScope. When new text input arrives, stale generation requests are dropped, ensuring that voice-cloned audio is generated only for the most recent turn. This prevents audio artifacts from overlapping requests during active conversations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →