# How the Mimi Codec Functions in Pocket‑TTS Voice Cloning

> Discover how the Mimi codec functions in Pocket-TTS voice cloning. Learn how it compresses audio into latent representations to match original speaker voices.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-09

---

**The Mimi codec compresses audio prompts into latent representations that preserve speaker-specific spectral characteristics, enabling Pocket‑TTS to generate new speech that matches the original voice.**

Voice cloning in Pocket‑TTS relies on a neural compression codec called **Mimi** to extract and reuse speaker identity from short audio samples. According to the kyutai-labs/pocket-tts source code, this codec bridges raw waveform processing and the flow‑LM's latent space through a multi-stage pipeline. Understanding how the Mimi codec functions in pocket-tts voice cloning reveals the technical foundation for zero-shot speaker adaptation.

## The Voice Cloning Pipeline

The Mimi codec operates as a bidirectional translator between audio waveforms and the flow language model's internal representation. When you provide a voice prompt, the system executes five distinct stages to capture and later reconstruct the speaker's timbre.

### Step 1: Encoding Raw Audio with SEANet

The process begins when the input waveform—whether from a `.wav` file or `.safetensors` tensor—passes through the **SEANet encoder**. This convolutional architecture produces a high-dimensional feature map that captures spectral details essential for speaker identity.

In [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) (lines 31-33), the `MimiModel` class initializes this via `self.encoder`, which instantiates `SEANetEncoder` from the modules layer. The encoder downsamples the raw audio into a compressed temporal representation while preserving acoustic features necessary for voice cloning.

### Step 2: Projecting to Flow‑Model Latent Space

Encoder outputs must align with the flow‑LM's expected frame rate and dimensionality. The `encoder_transformer`—a `ProjectedTransformer` defined in [`pocket_tts/modules/mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py) (lines 104-138)—projects these features into the model's latent space.

If the encoder's frame rate differs from the internal flow‑model rate, the `_to_framerate` method applies a 1-D convolutional down‑sampler (`ConvDownsample1d`) to reshape temporal resolution. This step occurs in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) (lines 33-55), ensuring the voice state matches the generation model's expectations.

### Step 3: Quantization (Optional)

The architecture includes a quantizer stage for additional compression, though the current implementation uses a **dummy quantizer** that passes latents through unchanged. The `DummyQuantizer` class in [`pocket_tts/modules/dummy_quantizer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/dummy_quantizer.py) serves as a placeholder that can be swapped for a learned neural quantizer without modifying the surrounding pipeline.

This quantizer is accessed via `MimiModel.quantizer` in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) (line 35), maintaining compatibility with future compression improvements.

### Step 4: Storing Voice State

Once processed, the latent tensor and internal transformer state (`mimi_state`) are packaged into a **model state dictionary**. This dictionary acts as a portable voice profile that can be cached, serialized, and reused across generation sessions.

The high-level API `get_state_for_audio_prompt` in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 408-416) orchestrates this process, returning a state object that conditions the flow‑LM on the specific speaker characteristics extracted from the prompt.

### Step 5: Decoding During Generation

When generating speech, the flow‑LM produces latent frames that must be converted back to audio. The `decode_from_latent` method in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) (lines 89-94) handles this by up‑sampling the latents (if needed via `ConvTrUpsample1d`) and feeding them to the **SEANet decoder**.

The decoder reconstructs the waveform while preserving the speaker timbre encoded in the original prompt, completing the voice cloning loop.

## Practical Implementation

The following example demonstrates the complete workflow: extracting a voice state from an audio file and generating cloned speech.

```python
from pathlib import Path
from pocket_tts import TTSModel

# Load the pre‑trained TTS model

model = TTSModel.load_model()

# Obtain a voice state from an audio prompt (local file, HF URL, or .safetensors)

voice_state = model.get_state_for_audio_prompt(
    Path("my_voice.wav")          # replace with your own wav file

    # or: "hf://kyutai/tts-voices/alba-mackenna/casual.wav"

)

# Generate speech that clones the prompt’s voice

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Hello, this is a cloned voice!",
    frames_after_eos=2,          # generate a few extra frames after EOS

)

# Save the result (requires soundfile or scipy)

import soundfile as sf
sf.write("cloned_output.wav", audio.squeeze().cpu().numpy(), samplerate=model.sample_rate)

```

For real-time applications, use the streaming interface to process audio chunks incrementally:

```python
for chunk in model.generate_audio_stream(voice_state, "Streaming example text…"):
    # chunk is a short tensor of audio samples; play or write immediately

    process_chunk(chunk)   # your custom handling

```

## Key Source Files and Components

The Mimi codec implementation spans several modules that handle specific aspects of the compression pipeline:

- **[`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)** – Implements `MimiModel`, orchestrating the SEANet encoder/decoder, transformers, and frame‑rate conversion logic.
- **[`pocket_tts/modules/mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py)** – Defines `ProjectedTransformer`, which maps encoder features into the flow‑LM latent space and vice‑versa.
- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)** – Provides the public `get_state_for_audio_prompt` API that invokes the Mimi codec and caches voice states.
- **[`pocket_tts/modules/seanet.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py)** – Contains `SEANetEncoder` and `SEANetDecoder`, the convolutional codec cores.
- **[`pocket_tts/modules/conv.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/conv.py)** – Supplies `ConvDownsample1d` and `ConvTrUpsample1d` for temporal resolution matching between codec and flow‑LM.
- **[`pocket_tts/modules/dummy_quantizer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/dummy_quantizer.py)** – Placeholder quantizer enabling future compression upgrades without pipeline changes.

## Summary

- **The Mimi codec compresses** audio prompts using a SEANet encoder to preserve speaker-specific spectral features.
- **Projection layers** align encoded features with the flow‑LM's latent space using transformers and convolutional resampling.
- **Voice states** are stored as model state dictionaries containing latent tensors and transformer states, enabling reuse across sessions.
- **Decoding** reconstructs waveforms via SEANet while maintaining the original speaker timbre, completing the voice cloning process.
- **The architecture** supports future quantization improvements through its modular dummy quantizer design.

## Frequently Asked Questions

### How does the Mimi codec preserve speaker identity during compression?

The codec uses a SEANet encoder that learns to retain spectral characteristics critical for speaker recognition while compressing the waveform. The subsequent transformer projects these features into a latent space that the flow language model can condition on, ensuring generated audio inherits the prompt's voice timbre.

### What audio formats can be used as prompts for voice cloning?

Pocket‑TTS accepts local `.wav` files, Hugging Face URLs (using the `hf://` protocol), or pre‑processed `.safetensors` files. The `get_state_for_audio_prompt` method in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) handles format detection and preprocessing automatically.

### Why does the Mimi codec use a dummy quantizer instead of a learned one?

The current implementation uses `DummyQuantizer` as a placeholder to maintain architectural compatibility. This design allows the development team to swap in a proper neural quantizer later for additional compression without changing the surrounding pipeline code or API signatures.

### Can the voice state be cached and reused across different text inputs?

Yes. The model state dictionary returned by `get_state_for_audio_prompt` can be serialized and reused with multiple `generate_audio` or `generate_audio_stream` calls. This enables efficient voice cloning where you extract the voice state once and generate hours of speech without re‑processing the original prompt.