# How to Use Custom Voice Cloning with Pocket TTS in Hugging Face Speech-to-Speech

> Learn how to use custom voice cloning with Pocket TTS in Hugging Face Speech-to-Speech. Clone voices from local files or Hugging Face repos effortlessly for your projects.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-11

---

**Pocket TTS supports real-time voice cloning from local audio files or Hugging Face repositories by passing a path or URL to `--pocket_tts_voice`, which extracts a voice state via `get_state_for_audio_prompt()` in the `PocketTTSHandler` class.**

The `huggingface/speech-to-speech` repository provides a modular pipeline for real-time speech-to-speech translation, featuring a **Pocket TTS** backend capable of cloning arbitrary voices from audio references. This voice cloning functionality is implemented in the `PocketTTSHandler` class, which extracts voice characteristics from presets, local files, or remote repositories using the `get_state_for_audio_prompt()` method. Understanding how to configure the `pocket_tts_voice` parameter enables you to generate personalized speech output tailored to specific speakers.

## Understanding the Pocket TTS Architecture

### The PocketTTSHandler Class

Located in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py), the `PocketTTSHandler` inherits from `BaseHandler[TTSIn, TTSOut]` and manages the voice cloning lifecycle. During initialization, the handler loads the Pocket TTS model onto the specified device and prepares the voice state that will be used for all subsequent text-to-speech generation.

### Voice State Extraction

The core cloning mechanism occurs in `get_state_for_audio_prompt()`, which processes the input specified by `pocket_tts_voice`. This method accepts preset names (such as `alba`, `marius`, `javert`, `jean`, `fantine`, `cosette`, `eponine`, or `azelma`), absolute paths to local WAV files, or Hugging Face repository URLs in the format `hf://kyutai/tts-voices/...`. The extracted voice state is cached for efficient streaming during the `process()` method, which yields `int16` audio blocks resampled from the model's native 24 kHz to the target sample rate using `scipy.signal.resample_poly`.

## Command-Line Voice Cloning Configuration

### Cloning from a Hugging Face Repository

To clone a voice hosted on the Hugging Face Hub, pass the repository URL prefixed with `hf://` to the `--pocket_tts_voice` argument. The handler automatically downloads the reference audio and metadata, extracting the voice state before generating speech.

```bash
python s2s_pipeline.py \
  --tts pocket \
  --pocket_tts_voice hf://kyutai/tts-voices/my_custom_voice \
  --pocket_tts_device cuda \
  --pocket_tts_sample_rate 16000

```

### Cloning from a Local Audio File

For private voice samples stored locally, provide the absolute path to the reference WAV file. The handler reads the waveform directly and computes the voice state without network requests.

```bash
python s2s_pipeline.py \
  --tts pocket \
  --pocket_tts_voice /absolute/path/to/my_reference.wav \
  --pocket_tts_device cpu

```

## Programmatic Voice Cloning in Python

For integration into custom applications, instantiate `PocketTTSHandler` directly with your voice source. The handler requires a `should_listen` event for pipeline coordination and accepts the same voice parameters as the CLI.

```python
from threading import Event
from speech_to_speech.TTS.pocket_tts_handler import PocketTTSHandler
from speech_to_speech.pipeline.handler_types import TTSIn

# Initialize the handler with custom voice

should_listen = Event()
handler = PocketTTSHandler(
    should_listen,
    device="cuda",
    voice="hf://kyutai/tts-voices/custom_voice",
    sample_rate=16000,
    blocksize=512,
)

# Setup loads the model and extracts voice state

handler.setup(
    should_listen,
    device="cuda",
    voice="hf://kyutai/tts-voices/custom_voice",
)

# Generate speech from text

tts_input = TTSIn(text="This is a cloned voice speaking.", language_code="en")
for audio_block in handler.process(tts_input):
    # audio_block is int16 numpy array of length blocksize

    play_audio(audio_block)  # Replace with your audio output logic

```

## Key Configuration Parameters

The following arguments control voice cloning behavior when using the Pocket TTS backend:

- **`pocket_tts_voice`**: Specifies the voice source. Accepts preset names (`jean`, `alba`, etc.), local file paths, or `hf://` URLs. Default: `jean`.
- **`pocket_tts_device`**: Compute device for model inference. Options: `cpu`, `cuda`, `mps`. Default: `cpu`.
- **`pocket_tts_sample_rate`**: Output audio sample rate. The model generates at 24 kHz internally, then resamples to this value. Default: `16000`.

## Summary

- **Pocket TTS voice cloning** is implemented in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) via the `PocketTTSHandler` class.
- Clone voices by setting `--pocket_tts_voice` to a preset name, local file path, or `hf://` repository URL.
- The `get_state_for_audio_prompt()` method extracts voice characteristics from the reference audio during handler setup.
- Audio output is automatically resampled from 24 kHz to your specified `pocket_tts_sample_rate` using `scipy.signal.resample_poly`.
- The handler integrates with the speculative turn tracker and cancel scope to ensure only fresh voice-cloned audio is generated.

## Frequently Asked Questions

### What audio file formats are supported for custom voice cloning?

The `PocketTTSHandler` reads reference audio using standard audio I/O libraries available in the environment. While the documentation emphasizes WAV files, any format supported by the underlying audio processing dependencies can be loaded from local paths or Hugging Face repositories. The reference audio should contain clear speech samples for optimal voice extraction.

### Can I switch between different cloned voices at runtime without restarting the pipeline?

Yes, you can change the voice by passing a different `--pocket_tts_voice` value when restarting the pipeline, or programmatically by creating a new handler instance with the desired voice parameter. The `setup()` method re-initializes the voice state via `get_state_for_audio_prompt()` for the new source, allowing dynamic switching between presets, local files, or remote repositories.

### How does the handler manage audio resampling for voice-cloned output?

Pocket TTS generates audio at a fixed 24 kHz sample rate internally. The handler calculates a resampling ratio based on your configured `pocket_tts_sample_rate` and applies `scipy.signal.resample_poly` to produce audio blocks matching the pipeline's default blocksize of 512 samples. This ensures compatible output regardless of the target sample rate.

### Does the Pocket TTS handler support concurrent voice cloning requests?

The handler inherits cancellation logic from `BaseHandler` that interacts with a `SpeculativeTurnTracker` and `CancelScope`. When new text input arrives, stale generation requests are dropped, ensuring that voice-cloned audio is generated only for the most recent turn. This prevents audio artifacts from overlapping requests during active conversations.