# How Pocket-TTS Achieves Voice Cloning with Audio Encoding: A Technical Deep Dive

> Discover how Pocket-TTS achieves voice cloning using audio encoding via Mimi neural codec and Flow-LM. Learn the technical details for efficient voice generation.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-09

---

**Pocket-TTS achieves voice cloning by encoding audio prompts into latent representations using the Mimi neural codec, projecting them into the Flow-LM conditioning space, and caching the resulting model state for reuse across generation tasks.**

The `kyutai-labs/pocket-tts` repository implements a zero-shot voice cloning system that converts audio prompts into reusable model states. By leveraging a neural audio codec and a Flow Language Model, the system extracts speaker characteristics from reference audio and applies them to new text inputs. This article examines how **voice cloning with audio encoding** works under the hood, tracing the specific implementation details found in the source code.

## The Voice Cloning Pipeline Architecture

The voice cloning process in Pocket-TTS transforms raw audio into a conditioning state through four distinct stages. Each stage is implemented in specific modules within the codebase.

### Loading and Preprocessing Audio Prompts

Voice cloning begins in `TTSModel.get_state_for_audio_prompt()` located in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 88-101). This method accepts multiple input formats: local file paths, HTTP/HuggingFace URLs, pre-encoded `.safetensors` files, or raw `torch.Tensor` objects.

If the input path ends with `.safetensors`, the state imports directly from disk. Otherwise, the method calls `audio_read` from [`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py) to load the waveform. The audio is then truncated to a maximum of 30 seconds and resampled to the model's native sample rate of 24 kHz using `convert_audio` from [`pocket_tts/data/audio_utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio_utils.py).

### Encoding Audio with the Mimi Neural Codec

Once preprocessed, the resampled waveform enters `_encode_audio()` in [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py) (lines 79-89), which delegates to `self.mimi.encode_to_latent()`.

The `MimiModel.encode_to_latent()` method in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) (lines 96-119) processes the signal through three sequential operations:

- **SEANetEncoder**: Extracts high-level acoustic features defined in [`pocket_tts/modules/seanet.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py)
- **Projected Transformer**: Applies attention mechanisms via [`pocket_tts/modules/mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py)
- **Downsampling**: Reduces to the global frame rate using `self._to_framerate`

The output is a latent tensor capturing the speaker's timbre, pitch, and prosodic characteristics.

### Projecting Latents into Flow-LM Space

The latent representation must align with the Flow Language Model's expected input dimensions. In [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py) (lines 86-89), the code transposes the latent tensor and applies a linear projection using `self.flow_lm.speaker_proj_weight`.

If the Flow-LM configuration requires a beginning-of-sequence token, the system concatenates `self.flow_lm.bos_before_voice` with the projected latent using `torch.cat([self.flow_lm.bos_before_voice, prompt], dim=1)` (lines 93-95).

### Creating the Reusable Model State

The final stage initializes the generation infrastructure. The `init_states()` method builds internal state dictionaries for both the Flow-LM and the Mimi decoder, pre-allocating KV caches for efficient incremental generation (lines 96-100).

The system then seeds these caches by running a single prompting step via `self._run_flow_lm_and_increment_step`, inserting the voice conditioning directly into the Flow-LM's hidden state. This state object can now generate unlimited text while maintaining the cloned voice characteristics.

## Performance Optimization via LRU Caching

To eliminate redundant computation when reusing the same voice prompt, Pocket-TTS wraps the public API with an LRU cache. The `_cached_get_state_for_audio_prompt` decorator (lines 81-86 in [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py)) ensures that subsequent calls with identical audio files return the pre-computed state instantly, bypassing the encode-and-project pipeline entirely.

## Implementation Details from Source Code

The voice cloning system relies on several interconnected components:

- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Core orchestration class implementing `get_state_for_audio_prompt`, audio encoding, and generation APIs
- **[`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)**: Neural audio codec with `encode_to_latent` and `decode_from_latent` methods
- **[`pocket_tts/modules/seanet.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py)**: SEANet encoder/decoder architecture for acoustic feature extraction
- **[`pocket_tts/modules/mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py)**: Projected transformer layers used within the Mimi codec
- **[`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py)**: Low-level WAV reading utilities (`audio_read`)
- **[`pocket_tts/data/audio_utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio_utils.py)**: Resampling helpers (`convert_audio`)

## Practical Code Examples

The following examples demonstrate the complete voice cloning workflow using the Pocket-TTS API.

Load the model and initialize the TTS engine:

```python
from pocket_tts import TTSModel

# Load model (CPU-compatible, works out-of-the-box)

model = TTSModel.load_model()

```

Create a voice state from a HuggingFace audio URL:

```python

# Clone voice from reference audio

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

```

Generate audio with the cloned voice:

```python

# Generate single utterance

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Hello from my cloned voice!",
    frames_after_eos=2,
    copy_state=True,
)

print(f"Generated audio shape: {audio.shape}")

```

Stream generation for real-time applications:

```python

# Stream audio chunks for immediate playback

for chunk in model.generate_audio_stream(voice_state, "Streaming example..."):
    # chunk is a 1-D tensor ready for audio sink

    play(chunk)  # e.g., using sounddevice

```

Reuse the same voice state efficiently:

```python

# Generate multiple sentences without re-encoding

sentences = ["First line.", "Second line.", "Third line."]
for s in sentences:
    audio = model.generate_audio(voice_state, s)
    save_wav(f"{s[:10]}.wav", audio, sr=model.sample_rate)

```

## Summary

- **Pocket-TTS** achieves voice cloning by extracting latent representations from audio prompts using the **Mimi neural codec**, then projecting these into the Flow-LM's conditioning space.
- The `TTSModel.get_state_for_audio_prompt()` method in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) handles the complete pipeline: loading, encoding, projecting, and state initialization.
- **SEANetEncoder** and **projected transformers** in the Mimi codec capture timbre, pitch, and prosody at 24 kHz.
- Generated **model states** include pre-allocated KV caches, enabling efficient reuse across multiple generation calls without re-encoding the reference audio.
- An **LRU cache** wrapper eliminates redundant computation when the same voice prompt is requested multiple times.

## Frequently Asked Questions

### What audio formats does Pocket-TTS accept for voice cloning?

Pocket-TTS accepts local file paths, HTTP/HuggingFace URLs, pre-encoded `.safetensors` files, or raw `torch.Tensor` objects. The `get_state_for_audio_prompt()` method automatically detects the format and routes it through `audio_read` for standard audio files or direct loading for `.safetensors` states.

### How does the Mimi codec encode speaker characteristics?

The `MimiModel.encode_to_latent()` method processes audio through a SEANetEncoder, projected transformer layers, and downsampling operations. This architecture captures high-fidelity representations of speaker timbre, pitch, and prosody while compressing the audio into a latent space suitable for the Flow Language Model.

### Can I reuse a voice state for multiple generation tasks?

Yes. Once created via `get_state_for_audio_prompt()`, the model state can be passed to `generate_audio()` or `generate_audio_stream()` repeatedly. The system employs an LRU cache to return pre-computed states instantly for identical audio prompts, making batch processing efficient.

### What is the maximum length for reference audio prompts?

The system optionally truncates input audio to 30 seconds during preprocessing. This occurs in the loading phase before resampling to 24 kHz, ensuring consistent encoding while managing memory usage and computational overhead.