# What Is the Mimi Neural Audio Codec and How It Is Used for Encoding and Decoding

> Discover the Mimi neural audio codec, a powerful system for compressing and reconstructing high-quality audio. Learn how it enables voice cloning and streaming synthesis in Pocket-TTS.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**The Mimi neural audio codec is a neural audio compression system that converts raw waveforms into compact latent representations and reconstructs high-quality audio, enabling voice cloning and streaming synthesis in Pocket-TTS.**

The **Mimi neural audio codec** serves as the core compression engine in the [kyutai-labs/pocket-tts](https://github.com/kyutai-labs/pocket-tts) repository. It bridges raw audio waveforms and the Flow-LM generative model by providing efficient encoding and decoding capabilities. This article examines the codec's architecture, public API, and integration within the text-to-speech pipeline based on the actual source implementation.

## Architecture of the Mimi Neural Audio Codec

The Mimi codec is composed of three distinct neural components that work together to compress and decompress audio. The architecture is implemented in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) within the `MimiModel` class.

| Component | Role | Source Location |
| --- | --- | --- |
| **SEANet Encoder / Decoder** | Convolutional audio codec operating at a high-level frame rate that extracts and reconstructs audio features while preserving fine-grained temporal detail. | [`SEANetEncoder`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py#L44) and [`SEANetDecoder`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py#L16) in [`pocket_tts/modules/seanet.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/seanet.py) |
| **ProjectedTransformer** | Lightweight transformer projecting encoder/decoder embeddings into the latent space used by the Flow-LM, aligning feature dimensions for bidirectional conversion. | [`ProjectedTransformer`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py#L104) in [`pocket_tts/modules/mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/mimi_transformer.py) |
| **DummyQuantizer** | Thin wrapper providing output projection (no true quantization in the inference-only version) that maps encoder output to the exact dimensionality expected by the Flow-LM. | [`DummyQuantizer`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/dummy_quantizer.py#L5) in [`pocket_tts/modules/dummy_quantizer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/dummy_quantizer.py) |

### SEANet Encoder and Decoder

The **SEANetEncoder** (defined at line 44 of [`seanet.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/seanet.py)) processes raw waveforms using convolutional layers to produce frame-level representations. The **SEANetDecoder** (line 16) performs the inverse operation, reconstructing waveforms from processed latent vectors. These components handle the high-level frame rate conversion while maintaining temporal fidelity.

### ProjectedTransformer

The **ProjectedTransformer** (line 104 of [`mimi_transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/mimi_transformer.py)) acts as a bridge between the SEANet components and the Flow-LM. It projects embeddings into the latent space required for generative modeling, enabling the codec to interface with the language model during both encoding and decoding phases.

### DummyQuantizer

The **DummyQuantizer** (line 5 of [`dummy_quantizer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/dummy_quantizer.py)) provides dimensionality matching without performing actual quantization. This design choice reflects the inference-only nature of the current implementation, where the codec operates as a continuous compression system rather than a quantized discrete representation.

## Encoding and Decoding API

The `MimiModel` class exposes two primary methods for audio compression and reconstruction:

- **`encode_to_latent(x: torch.Tensor) -> torch.Tensor`**: Accepts a batch of waveforms with shape `[B, C, T]` (batch, channels, samples), pads them to an integer multiple of the frame size, and processes them through the encoder, encoder-side transformer, and downsampling layers. The output is an unquantized latent tensor ready for the Flow-LM.

- **`decode_from_latent(latent: torch.Tensor, mimi_state) -> torch.Tensor`**: Receives a latent tensor produced by the Flow-LM, upsamples it to the encoder frame rate, applies the decoder-side transformer, and passes it through the SEANet decoder to obtain a reconstructed waveform.

Both methods are wrapped with timing utilities (`display_execution_time`) to monitor performance characteristics during execution.

## Integration in the TTS Pipeline

The Mimi codec integrates into the Pocket-TTS pipeline at two critical stages:

1. **Audio Prompt Encoding**: When processing voice cloning prompts, `TTSModel.get_state_for_audio_prompt` calls `_encode_audio`, which internally invokes `self.mimi.encode_to_latent`. The resulting latent tensor is linearly projected into the Flow-LM's latent space using `speaker_proj_weight`.

2. **Latent Decoding**: During generation, the Flow-LM produces latent vectors that are fed to `MimiModel.decode_from_latent`. This restores the waveform chunk-by-chunk, enabling streaming playback without waiting for complete file generation.

This architecture allows the system to operate efficiently on CPU-only environments while maintaining real-time streaming capabilities.

## Practical Code Examples

### Encoding and Decoding Audio

The following example demonstrates standalone encoding and decoding using the Mimi codec:

```python
import torch
from pocket_tts import TTSModel

# Load a pre-trained model (weights are downloaded automatically)

model = TTSModel.load_model()

# Encode a raw audio tensor to Mimi latent space

# `audio` must be a FloatTensor of shape [channels, samples]

# Example: a 1-second clip at 24 kHz -> shape [1, 24000]

audio = torch.randn(1, 24000)  # replace with real audio data

latent = model.mimi.encode_to_latent(audio.unsqueeze(0))  # batch dim added

print("Latent shape:", latent.shape)  # e.g. [1, 80, 250] (B, latent_dim, frames)

# Decode the latent back into a waveform

# Initialise a fresh Mimi state for decoding

mimi_state = model.mimi.init_states(batch_size=1, sequence_length=latent.shape[-1])

# Decode

reconstructed = model.mimi.decode_from_latent(latent, mimi_state)

print("Reconstructed audio shape:", reconstructed.shape)  # [1, 1, samples]

```

### Full Voice Cloning Workflow

This example shows the complete pipeline from audio prompt to synthesized speech:

```python

# Load an audio prompt (e.g., a short speaker sample)

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Generate speech for a given text, streaming chunks

for chunk in model.generate_audio_stream(voice_state, "Hello, world!"):
    # `chunk` is a 1-D Tensor of raw samples ready for playback or saving

    # (e.g., write to a WAV file with scipy.io.wavfile.write)

    pass

```

## Summary

- The **Mimi neural audio codec** provides neural compression for raw audio waveforms in the Pocket-TTS system.
- It consists of **SEANet** convolutional encoder/decoder pairs, a **ProjectedTransformer** for latent space alignment, and a **DummyQuantizer** for dimensionality projection.
- The `encode_to_latent` method in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) compresses waveforms for the Flow-LM, while `decode_from_latent` reconstructs audio for streaming output.
- The codec enables voice cloning by encoding audio prompts and supports real-time generation through chunk-based decoding.

## Frequently Asked Questions

### How does the Mimi neural audio codec differ from traditional audio codecs?

Traditional audio codecs rely on hand-designed algorithms like psychoacoustic models, whereas the Mimi neural audio codec uses learned convolutional representations through the SEANet architecture. This neural approach allows the system to optimize compression specifically for the downstream Flow-LM generative model rather than human perceptual limits alone.

### What is the purpose of the DummyQuantizer in the codec architecture?

The **DummyQuantizer** serves as a dimensionality adapter that maps the encoder's output to the exact latent dimensionality expected by the Flow-LM. According to the source code in [`pocket_tts/modules/dummy_quantizer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/dummy_quantizer.py), it performs no true quantization in the inference-only version, operating instead as a learned projection layer that maintains continuous representations.

### What audio format does the Mimi codec expect for encoding?

The `encode_to_latent` method expects a **FloatTensor** of shape `[B, C, T]` representing batch, channels, and time samples. The code examples demonstrate 24 kHz sampling rates, though the architecture handles various lengths by padding inputs to integer multiples of the frame size before processing through the SEANet encoder.

### Is the Mimi codec suitable for real-time streaming applications?

Yes, the codec supports streaming through its `decode_from_latent` method, which processes audio chunk-by-chunk rather than requiring full sequences. The implementation includes `display_execution_time` wrappers for performance monitoring, and the architecture is optimized for CPU-only operation, making it suitable for low-latency synthesis scenarios.