# How to Implement Custom TTS Voice Cloning with Pocket TTS in the Speech-to-Speech Pipeline

> Learn to implement custom TTS voice cloning with Pocket TTS using Hugging Face speech-to-speech. Easily set voice embeddings from presets, local files, or repo URLs.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-30

---

**You can clone custom voices in the Hugging Face speech-to-speech pipeline by setting the `pocket_tts_voice` argument to a preset name, local audio file path, or Hugging Face repository URL, which the `PocketTTSHandler` converts into a voice embedding via `get_state_for_audio_prompt`.**

The huggingface/speech-to-speech repository integrates Pocket TTS (Kyutai Labs' open-source model) to enable real-time text-to-speech synthesis with custom voice cloning capabilities. By configuring the `PocketTTSHandlerArguments`, you can specify voice sources ranging from built-in presets to custom audio files or Hugging Face hosted embeddings. This guide covers how to implement custom TTS voice cloning with Pocket TTS using both command-line interfaces and programmatic Python configuration.

## Understanding Pocket TTS Architecture

The Pocket TTS integration centers around the `PocketTTSHandler` class defined in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py). When initializing the pipeline with `module_kwargs.tts = "pocket"`, the factory function `get_tts_handler` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) instantiates this handler using parameters from `PocketTTSHandlerArguments` located in [`src/speech_to_speech/arguments_classes/pocket_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/pocket_tts_arguments.py).

The handler manages model initialization, voice state loading, and audio streaming. It automatically resamples the native 24 kHz output to your target sample rate while processing audio in configurable blocks for real-time playback.

## Voice Source Options for Custom Cloning

The `pocket_tts_voice` parameter accepts three distinct input types, enabling flexible voice cloning workflows:

- **Preset Names** — Built-in voices shipped with the model such as `"alba"` or `"jean"` (the default). These require no additional files and load immediately.
- **Local Audio Files** — Absolute paths to WAV files on your local filesystem, such as `"/path/to/my_voice.wav"`. The handler extracts voice embeddings directly from these files for instant cloning.
- **Hugging Face Repository Paths** — References to voice embeddings stored in Hugging Face repositories using the `hf://` protocol, like `"hf://kyutai/tts-voices/my_voice"`. This enables sharing and versioning custom voices across teams.

## Command-Line Voice Configuration

Configure custom voice cloning via CLI by passing the `--pocket_tts_voice` argument when launching the pipeline.

Use a built-in preset:

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice alba

```

Clone from a local audio file:

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice /home/user/my_voice.wav

```

Load a voice from a Hugging Face repository:

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice "hf://kyutai/tts-voices/custom_voice"

```

## Programmatic Implementation

For custom pipelines, instantiate `PocketTTSHandlerArguments` and pass it to the TTS handler factory:

```python
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.s2s_pipeline import get_tts_handler
from threading import Event
from queue import Queue

# Configure voice cloning parameters

pocket_args = PocketTTSHandlerArguments(
    pocket_tts_device="cpu",  # or "cuda"/"mps"

    pocket_tts_voice="hf://kyutai/tts-voices/custom_voice",
    pocket_tts_sample_rate=16000,
    pocket_tts_blocksize=512,
    pocket_tts_max_tokens=80,
)

# Initialize pipeline queues and events

stop_event = Event()
lm_response_q = Queue()
audio_out_q = Queue()
should_listen = Event()

# Create the handler with custom voice settings

tts_handler = get_tts_handler(
    module_kwargs=type("M", (), {"tts": "pocket"}),
    stop_event=stop_event,
    lm_response_queue=lm_response_q,
    send_audio_chunks_queue=audio_out_q,
    should_listen=should_listen,
    chat_tts_handler_kwargs=None,
    facebook_mms_tts_handler_kwargs=None,
    pocket_tts_handler_kwargs=pocket_args,
    kokoro_tts_handler_kwargs=None,
    qwen3_tts_handler_kwargs=None,
)

```

## Internal Voice Loading Mechanism

Inside [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py), the `setup()` method processes the `voice` parameter through the Pocket TTS library's abstraction layer:

```python

# From PocketTTSHandler.setup(...)

logger.info(f"Loading voice: {voice}")
self.voice_state = self.model.get_state_for_audio_prompt(voice)

```

The `get_state_for_audio_prompt` method handles all three voice source types transparently, whether you're using a preset string, local file path, or Hugging Face URL. This unified interface simplifies voice management while supporting real-time streaming requirements.

## Audio Processing and Streaming Configuration

The handler automatically manages audio format conversions and streaming buffers. Key parameters include:

- **Sample Rate** — Defaults to 16 kHz, with automatic resampling from the model's native 24 kHz output
- **Blocksize** — Configurable chunk size (default 512 samples) controlling the granularity of audio streaming to downstream components
- **Device** — CPU, CUDA, or MPS selection for model inference

These settings are defined in `PocketTTSHandlerArguments` and apply consistently across both CLI and programmatic usage.

## Summary

- **PocketTTSHandler** in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) manages the complete voice cloning pipeline from loading to streaming
- **Three voice source types** are supported: preset names (`"jean"`, `"alba"`), local WAV files, and Hugging Face repository paths (`hf://`)
- **Configuration** occurs through `PocketTTSHandlerArguments` in [`src/speech_to_speech/arguments_classes/pocket_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/pocket_tts_arguments.py)
- **Voice embedding extraction** happens via `get_state_for_audio_prompt` transparently across all source types
- **Real-time streaming** uses configurable block sizes with automatic resampling from 24 kHz to your target sample rate

## Frequently Asked Questions

### What audio file formats does Pocket TTS accept for voice cloning?

Pocket TTS accepts standard WAV files for local voice cloning. When providing a file path to the `pocket_tts_voice` argument, ensure the file is a valid WAV format that the `get_state_for_audio_prompt` method can process into voice embeddings.

### Can I use Pocket TTS voice cloning with GPU acceleration?

Yes. Set `pocket_tts_device="cuda"` in the `PocketTTSHandlerArguments` or pass the appropriate device flag via CLI. The handler supports CPU, CUDA, and MPS backends for model inference according to the source code implementation.

### How do I share custom voices across different deployments?

Store your voice embeddings in a Hugging Face repository and reference them using the `hf://` protocol format (e.g., `hf://kyutai/tts-voices/custom_voice`). This allows consistent voice cloning across multiple environments without transferring local audio files.

### What is the default voice if I don't specify a custom voice?

If no voice is specified, Pocket TTS defaults to the `"jean"` preset voice, as defined in the handler's argument dataclass and verified in the test suite at [`tests/test_cli_defaults.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_cli_defaults.py).