# How Voice Cloning Works in pocket‑tts Using `get_state_for_audio_prompt()`

> Discover how pocket-tts voice cloning uses get_state_for_audio_prompt to extract a compact VoiceState from audio for efficient streaming generation.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**Voice cloning in `pocket‑tts` works by extracting a compact `VoiceState` from a reference audio file via `get_state_for_audio_prompt()`, which encodes the audio through the Mimi neural codec and caches the result for efficient reuse during streaming generation.**

The `pocket‑tts` repository by Kyutai Labs implements zero-shot voice cloning through a conditioning mechanism that leverages the Mimi audio codec and a Flow-based Language Model (Flow-LM). At the heart of this system is the `get_state_for_audio_prompt()` method defined in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), which transforms raw audio into a structured voice state that persists throughout the text-to-speech generation process.

## The Voice Cloning Pipeline

The `get_state_for_audio_prompt()` method orchestrates a five-stage pipeline to convert a reference waveform into a conditioning signal. Each stage is implemented across specific modules in the codebase.

### Step 1: Audio Loading and Preprocessing

The process begins in [`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py), where the input waveform undergoes standardization before neural encoding. The system reads the raw audio file, resamples it to the model's internal **24 kHz** sample rate, and applies normalization to ensure consistent amplitude levels.

- **`audio.read_wav`**: Loads the raw bytes from the file path.
- **`audio.resample`**: Converts the waveform to 24 kHz to match the Mimi encoder's expectations.

### Step 2: Mimi Codec Encoding

The preprocessed audio is fed into the **Mimi encoder** located in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py). The `MimiModel.encode()` function performs a dual extraction:

- **`latent_codes`**: A sequence of discrete tokens capturing acoustic content and temporal structure.
- **`speaker_embedding`**: A continuous vector representing the speaker's unique timbre and vocal characteristics.

This compression is critical because it distills the reference audio into a compact representation that the generative model can efficiently process.

### Step 3: VoiceState Construction

The extracted `latent_codes` and `speaker_embedding` are packaged into a **`VoiceState`** dataclass defined within [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py). This immutable structure serves as the standardized interface between the cloning frontend and the generation backend, ensuring that all acoustic conditioning information travels together through the pipeline.

### Step 4: LRU Caching for Performance

Voice cloning is computationally expensive due to the Mimi encoder's forward pass. To optimize repeated use of the same reference voice, the method is wrapped with **`@functools.lru_cache(maxsize=64)`** on an internal helper called `_cached_get_state_for_audio_prompt`. When you call `get_state_for_audio_prompt()` with the same file path, the system returns the cached `VoiceState` instantly, eliminating redundant encoding overhead.

### Step 5: Conditioning the Flow-LM

During generation, `generate_audio_stream()` accepts the `VoiceState` object via the `voice_state` parameter. The method injects these embeddings into the Flow-LM's internal state at every streaming step, ensuring the synthesized audio maintains consistent speaker identity with the reference prompt throughout the entire utterance.

## Practical Implementation with `get_state_for_audio_prompt()`

The public API abstracts the complexity of the pipeline behind a single method call. Below are practical patterns for implementing voice cloning in your applications.

### Basic Voice Cloning

This example demonstrates the complete workflow from loading a reference voice to generating cloned speech:

```python
from pocket_tts import TTSModel

# Initialize the model (downloads weights on first run)

tts = TTSModel()

# Path to your reference audio (WAV format, any length)

voice_path = "examples/voice_prompt.wav"

# Extract the voice state - this performs the cloning

voice_state = tts.get_state_for_audio_prompt(voice_path)

# Generate speech using the cloned voice identity

text = "Hello, this is a cloned voice speaking."
audio_chunks = tts.generate_audio_stream(
    text,
    voice_state=voice_state,  # Condition on the cloned voice

    temperature=0.7,
)

# Save the streamed output

with open("output.wav", "wb") as f:
    for chunk in audio_chunks:
        f.write(chunk)

```

### Reusing Cached Voice States

Because `get_state_for_audio_prompt()` implements LRU caching, subsequent calls with the same path return instantly:

```python
from pocket_tts import TTSModel

tts = TTSModel()
voice_path = "examples/voice_prompt.wav"

# First call triggers Mimi encoding (slow)

base_voice = tts.get_state_for_audio_prompt(voice_path)

# Subsequent calls hit the cache (instant)

for sentence in ["First sentence.", "Second sentence.", "Third sentence."]:
    audio = tts.generate_audio_stream(
        sentence,
        voice_state=tts.get_state_for_audio_prompt(voice_path),
    )
    # Process audio chunks...

```

## Key Source Files and Architecture

Understanding the repository structure helps when customizing or debugging the voice cloning pipeline:

- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Contains `TTSModel`, `get_state_for_audio_prompt()`, the `VoiceState` dataclass, and the LRU cache wrapper logic.
- **[`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)**: Implements the Mimi encoder/decoder that extracts latent codes and speaker embeddings from raw audio.
- **[`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py)**: Provides utility functions for reading, resampling, and normalizing WAV files before they enter the encoder.
- **[`pocket_tts/__init__.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/__init__.py)**: Exposes the public `TTSModel` API for direct imports.

## Summary

- **`get_state_for_audio_prompt()`** serves as the entry point for voice cloning, converting reference audio into a reusable `VoiceState` object.
- The method relies on the **Mimi codec** ([`mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/mimi.py)) to extract discrete latent tokens and continuous speaker embeddings from the input waveform.
- **LRU caching** (`maxsize=64`) prevents redundant encoding of identical voice prompts, significantly improving performance for batch generation.
- The resulting `VoiceState` conditions the **Flow-LM** during `generate_audio_stream()`, ensuring consistent speaker characteristics throughout synthesis.
- Audio preprocessing happens in [`audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/audio.py), which handles resampling to **24 kHz** and normalization.

## Frequently Asked Questions

### What is the purpose of `get_state_for_audio_prompt()`?

The method extracts a structured representation of a reference voice that the TTS model can use as a conditioning signal. It encapsulates the entire preprocessing and encoding pipeline, returning a `VoiceState` object that contains both the acoustic content (latent codes) and speaker identity (embedding) needed to guide the Flow-LM during generation.

### How does the Mimi codec contribute to voice cloning?

The Mimi codec ([`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)) functions as a neural audio compressor. Its encoder produces two distinct outputs: discrete latent codes that capture the temporal acoustic structure, and a continuous speaker embedding that encodes timbre. Together, these allow the generative model to reproduce both the speaking style and the specific vocal characteristics of the reference audio.

### Why is the voice state cached with `lru_cache`?

Encoding a voice prompt requires a full forward pass through the Mimi neural network, which involves significant computation. By wrapping the internal `_cached_get_state_for_audio_prompt()` with `@functools.lru_cache(maxsize=64)`, the system stores the resulting `VoiceState` in memory. Subsequent requests for the same file path return the cached object instantly, reducing latency when generating multiple utterances with the same cloned voice.

### Can I use multiple different voice prompts in the same session?

Yes. You can call `get_state_for_audio_prompt()` with different file paths to create distinct `VoiceState` objects, and the cache will retain them up to the 64-entry limit. Pass the appropriate `voice_state` to each `generate_audio_stream()` call to switch between different cloned voices or the default speaker at runtime.