What Is the Mimi Neural Audio Codec and How It Is Used for Encoding and Decoding

The Mimi neural audio codec is a neural audio compression system that converts raw waveforms into compact latent representations and reconstructs high-quality audio, enabling voice cloning and streaming synthesis in Pocket-TTS.

The Mimi neural audio codec serves as the core compression engine in the kyutai-labs/pocket-tts repository. It bridges raw audio waveforms and the Flow-LM generative model by providing efficient encoding and decoding capabilities. This article examines the codec's architecture, public API, and integration within the text-to-speech pipeline based on the actual source implementation.

Architecture of the Mimi Neural Audio Codec

The Mimi codec is composed of three distinct neural components that work together to compress and decompress audio. The architecture is implemented in pocket_tts/models/mimi.py within the MimiModel class.

Component Role Source Location
SEANet Encoder / Decoder Convolutional audio codec operating at a high-level frame rate that extracts and reconstructs audio features while preserving fine-grained temporal detail. SEANetEncoder and SEANetDecoder in pocket_tts/modules/seanet.py
ProjectedTransformer Lightweight transformer projecting encoder/decoder embeddings into the latent space used by the Flow-LM, aligning feature dimensions for bidirectional conversion. ProjectedTransformer in pocket_tts/modules/mimi_transformer.py
DummyQuantizer Thin wrapper providing output projection (no true quantization in the inference-only version) that maps encoder output to the exact dimensionality expected by the Flow-LM. DummyQuantizer in pocket_tts/modules/dummy_quantizer.py

SEANet Encoder and Decoder

The SEANetEncoder (defined at line 44 of seanet.py) processes raw waveforms using convolutional layers to produce frame-level representations. The SEANetDecoder (line 16) performs the inverse operation, reconstructing waveforms from processed latent vectors. These components handle the high-level frame rate conversion while maintaining temporal fidelity.

ProjectedTransformer

The ProjectedTransformer (line 104 of mimi_transformer.py) acts as a bridge between the SEANet components and the Flow-LM. It projects embeddings into the latent space required for generative modeling, enabling the codec to interface with the language model during both encoding and decoding phases.

DummyQuantizer

The DummyQuantizer (line 5 of dummy_quantizer.py) provides dimensionality matching without performing actual quantization. This design choice reflects the inference-only nature of the current implementation, where the codec operates as a continuous compression system rather than a quantized discrete representation.

Encoding and Decoding API

The MimiModel class exposes two primary methods for audio compression and reconstruction:

  • encode_to_latent(x: torch.Tensor) -> torch.Tensor: Accepts a batch of waveforms with shape [B, C, T] (batch, channels, samples), pads them to an integer multiple of the frame size, and processes them through the encoder, encoder-side transformer, and downsampling layers. The output is an unquantized latent tensor ready for the Flow-LM.

  • decode_from_latent(latent: torch.Tensor, mimi_state) -> torch.Tensor: Receives a latent tensor produced by the Flow-LM, upsamples it to the encoder frame rate, applies the decoder-side transformer, and passes it through the SEANet decoder to obtain a reconstructed waveform.

Both methods are wrapped with timing utilities (display_execution_time) to monitor performance characteristics during execution.

Integration in the TTS Pipeline

The Mimi codec integrates into the Pocket-TTS pipeline at two critical stages:

  1. Audio Prompt Encoding: When processing voice cloning prompts, TTSModel.get_state_for_audio_prompt calls _encode_audio, which internally invokes self.mimi.encode_to_latent. The resulting latent tensor is linearly projected into the Flow-LM's latent space using speaker_proj_weight.

  2. Latent Decoding: During generation, the Flow-LM produces latent vectors that are fed to MimiModel.decode_from_latent. This restores the waveform chunk-by-chunk, enabling streaming playback without waiting for complete file generation.

This architecture allows the system to operate efficiently on CPU-only environments while maintaining real-time streaming capabilities.

Practical Code Examples

Encoding and Decoding Audio

The following example demonstrates standalone encoding and decoding using the Mimi codec:

import torch
from pocket_tts import TTSModel

# Load a pre-trained model (weights are downloaded automatically)

model = TTSModel.load_model()

# Encode a raw audio tensor to Mimi latent space

# `audio` must be a FloatTensor of shape [channels, samples]

# Example: a 1-second clip at 24 kHz -> shape [1, 24000]

audio = torch.randn(1, 24000)  # replace with real audio data

latent = model.mimi.encode_to_latent(audio.unsqueeze(0))  # batch dim added

print("Latent shape:", latent.shape)  # e.g. [1, 80, 250] (B, latent_dim, frames)

# Decode the latent back into a waveform

# Initialise a fresh Mimi state for decoding

mimi_state = model.mimi.init_states(batch_size=1, sequence_length=latent.shape[-1])

# Decode

reconstructed = model.mimi.decode_from_latent(latent, mimi_state)

print("Reconstructed audio shape:", reconstructed.shape)  # [1, 1, samples]

Full Voice Cloning Workflow

This example shows the complete pipeline from audio prompt to synthesized speech:


# Load an audio prompt (e.g., a short speaker sample)

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Generate speech for a given text, streaming chunks

for chunk in model.generate_audio_stream(voice_state, "Hello, world!"):
    # `chunk` is a 1-D Tensor of raw samples ready for playback or saving

    # (e.g., write to a WAV file with scipy.io.wavfile.write)

    pass

Summary

  • The Mimi neural audio codec provides neural compression for raw audio waveforms in the Pocket-TTS system.
  • It consists of SEANet convolutional encoder/decoder pairs, a ProjectedTransformer for latent space alignment, and a DummyQuantizer for dimensionality projection.
  • The encode_to_latent method in pocket_tts/models/mimi.py compresses waveforms for the Flow-LM, while decode_from_latent reconstructs audio for streaming output.
  • The codec enables voice cloning by encoding audio prompts and supports real-time generation through chunk-based decoding.

Frequently Asked Questions

How does the Mimi neural audio codec differ from traditional audio codecs?

Traditional audio codecs rely on hand-designed algorithms like psychoacoustic models, whereas the Mimi neural audio codec uses learned convolutional representations through the SEANet architecture. This neural approach allows the system to optimize compression specifically for the downstream Flow-LM generative model rather than human perceptual limits alone.

What is the purpose of the DummyQuantizer in the codec architecture?

The DummyQuantizer serves as a dimensionality adapter that maps the encoder's output to the exact latent dimensionality expected by the Flow-LM. According to the source code in pocket_tts/modules/dummy_quantizer.py, it performs no true quantization in the inference-only version, operating instead as a learned projection layer that maintains continuous representations.

What audio format does the Mimi codec expect for encoding?

The encode_to_latent method expects a FloatTensor of shape [B, C, T] representing batch, channels, and time samples. The code examples demonstrate 24 kHz sampling rates, though the architecture handles various lengths by padding inputs to integer multiples of the frame size before processing through the SEANet encoder.

Is the Mimi codec suitable for real-time streaming applications?

Yes, the codec supports streaming through its decode_from_latent method, which processes audio chunk-by-chunk rather than requiring full sequences. The implementation includes display_execution_time wrappers for performance monitoring, and the architecture is optimized for CPU-only operation, making it suitable for low-latency synthesis scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →