# How to Configure Custom Voice Cloning with Pocket TTS in the Speech-to-Speech Pipeline

> Learn how to configure custom voice cloning with Pocket TTS in the speech-to-speech pipeline. Easily set preset names, local files, or repo URLs for unique voice generation.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-10

---

**You can configure custom voice cloning with Pocket TTS by setting the `pocket_tts_voice` argument to a preset name, local audio file path, or Hugging Face repository URL, which the `PocketTTSHandler` processes through the `get_state_for_audio_prompt` method.**

The `huggingface/speech-to-speech` repository integrates Kyutai Labs' Pocket TTS model through a dedicated handler that supports zero-shot voice cloning. When you select Pocket TTS as your text-to-speech backend, the pipeline instantiates `PocketTTSHandler` from [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) using configuration parameters defined in `PocketTTSHandlerArguments`.

## Understanding Pocket TTS Integration

The Pocket TTS integration follows the modular architecture of the speech-to-speech pipeline. When `module_kwargs.tts` is set to `"pocket"`, the factory function `get_tts_handler` in [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) initializes the handler with your custom arguments.

The handler manages the entire lifecycle of voice synthesis: loading the Kyutai model onto your specified device, preparing the voice embedding, generating audio at 24 kHz, and resampling to your target sample rate for real-time streaming.

## Voice Configuration Options

The `voice` parameter in `PocketTTSHandlerArguments` accepts three distinct input types, giving you flexibility in how you source your cloning target.

### Preset Voices

Pocket TTS ships with built-in voice presets including `"alba"` and `"jean"`. The default voice is `"jean"` as specified in the handler's docstring and argument defaults. These require no external files and load immediately from the model weights.

### Local Audio Files

You can clone any speaker by providing a path to a local WAV file. The handler extracts a voice embedding from this audio and uses it to condition the generation. The source audio should be clear speech for optimal cloning quality.

### Hugging Face Repository Paths

For shared or version-controlled voices, pass a Hugging Face repository path using the `hf://` protocol. The handler downloads the voice embedding from the specified repository location, enabling team collaboration and reproducible deployments.

## Configuring Voice Cloning via Command Line

The speech-to-speech CLI exposes the `pocket_tts_voice` argument to configure your desired voice source. Here are the three supported patterns:

**Use a built-in preset voice:**

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice alba

```

**Clone a custom voice from a local audio file:**

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice /home/user/my_voice.wav

```

**Load a voice from a Hugging Face repository:**

```bash
python -m speech_to_speech.main --tts pocket \
    --pocket_tts_voice "hf://kyutai/tts-voices/custom_voice"

```

## Configuring Voice Cloning Programmatically

For custom pipelines or research applications, import `PocketTTSHandlerArguments` from [`src/speech_to_speech/arguments_classes/pocket_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/pocket_tts_arguments.py) and instantiate the handler directly.

```python
from speech_to_speech.arguments_classes.pocket_tts_arguments import PocketTTSHandlerArguments
from speech_to_speech.s2s_pipeline import get_tts_handler
from queue import Queue
from threading import Event

# Configure the voice and processing parameters

pocket_args = PocketTTSHandlerArguments(
    pocket_tts_device="cuda",               # or "cpu", "mps"

    pocket_tts_voice="hf://kyutai/tts-voices/custom_voice",
    pocket_tts_sample_rate=16000,
    pocket_tts_blocksize=512,
    pocket_tts_max_tokens=80,
)

# Create required pipeline components

stop_event = Event()
lm_response_q = Queue()
audio_out_q = Queue()
should_listen = Event()

# Instantiate the handler

tts_handler = get_tts_handler(
    module_kwargs=type("M", (), {"tts": "pocket"}),
    stop_event=stop_event,
    lm_response_queue=lm_response_q,
    send_audio_chunks_queue=audio_out_q,
    should_listen=should_listen,
    chat_tts_handler_kwargs=None,
    facebook_mms_tts_handler_kwargs=None,
    pocket_tts_handler_kwargs=pocket_args,
    kokoro_tts_handler_kwargs=None,
    qwen3_tts_handler_kwargs=None,
)

```

## Technical Implementation Details

The voice cloning mechanism centers on the `setup()` method in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py). When the handler initializes, it processes the `voice` argument through the following code:

```python

# Inside PocketTTSHandler.setup(...)

logger.info(f"Loading voice: {voice}")
self.voice_state = self.model.get_state_for_audio_prompt(voice)

```

The `get_state_for_audio_prompt` method (provided by the underlying `pocket_tts` library) abstracts the loading logic for all three voice source types. It returns a voice state tensor that conditions the autoregressive generation throughout the streaming process.

**Audio Processing Pipeline:**

- **Native generation:** The model generates audio at 24 kHz
- **Resampling:** The handler automatically resamples to your configured `sample_rate` (default 16 kHz)
- **Streaming:** Audio emits in blocks of `blocksize` samples (default 512) to minimize latency for real-time playback

## Summary

- **Custom voice cloning with Pocket TTS** requires setting the `pocket_tts_voice` argument to either a preset name, local WAV path, or Hugging Face repository URL
- The `PocketTTSHandler` class in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) manages voice loading through `get_state_for_audio_prompt` and handles resampling from 24 kHz to your target sample rate
- Configuration is controlled via `PocketTTSHandlerArguments` in [`src/speech_to_speech/arguments_classes/pocket_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/pocket_tts_arguments.py), exposing parameters for device, sample rate, blocksize, and max tokens
- The default voice is `"jean"` when no custom voice is specified
- The handler streams audio in 512-sample blocks by default, enabling real-time speech-to-speech conversation

## Frequently Asked Questions

### What voice formats does Pocket TTS support for cloning?

Pocket TTS accepts three voice identifier formats: built-in string presets like `"alba"` or `"jean"`, absolute file paths to local WAV files (e.g., `/path/to/voice.wav`), and Hugging Face repository URLs using the `hf://` protocol (e.g., `hf://kyutai/tts-voices/custom_voice`). The handler automatically detects the format and processes it through `get_state_for_audio_prompt`.

### How do I change the audio sample rate in Pocket TTS?

Set the `pocket_tts_sample_rate` argument when configuring `PocketTTSHandlerArguments`. While the model natively generates at 24 kHz, the handler resamples to your specified rate (default 16 kHz) before queuing audio for playback. This parameter is available in both CLI (`--pocket_tts_sample_rate`) and programmatic interfaces.

### Can I use Pocket TTS with GPU acceleration?

Yes. Set `pocket_tts_device="cuda"` in your `PocketTTSHandlerArguments` or pass `--pocket_tts_device cuda` via CLI. The handler calls `.to(device)` on the model during initialization in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py), supporting CUDA, MPS (Apple Silicon), and CPU backends.

### Where is the default voice configured in the source code?

The default voice `"jean"` is defined in [`src/speech_to_speech/arguments_classes/pocket_tts_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/pocket_tts_arguments.py) within the `PocketTTSHandlerArguments` dataclass. The test suite in [`tests/test_cli_defaults.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/test_cli_defaults.py) verifies that this default is correctly wired into the pipeline when no custom voice is specified.