What Is the Mimi Neural Audio Codec and How It Is Used for Encoding and Decoding
The Mimi neural audio codec is a neural audio compression system that converts raw waveforms into compact latent representations and reconstructs high-quality audio, enabling voice cloning and streaming synthesis in Pocket-TTS.
The Mimi neural audio codec serves as the core compression engine in the kyutai-labs/pocket-tts repository. It bridges raw audio waveforms and the Flow-LM generative model by providing efficient encoding and decoding capabilities. This article examines the codec's architecture, public API, and integration within the text-to-speech pipeline based on the actual source implementation.
Architecture of the Mimi Neural Audio Codec
The Mimi codec is composed of three distinct neural components that work together to compress and decompress audio. The architecture is implemented in pocket_tts/models/mimi.py within the MimiModel class.
| Component | Role | Source Location |
|---|---|---|
| SEANet Encoder / Decoder | Convolutional audio codec operating at a high-level frame rate that extracts and reconstructs audio features while preserving fine-grained temporal detail. | SEANetEncoder and SEANetDecoder in pocket_tts/modules/seanet.py |
| ProjectedTransformer | Lightweight transformer projecting encoder/decoder embeddings into the latent space used by the Flow-LM, aligning feature dimensions for bidirectional conversion. | ProjectedTransformer in pocket_tts/modules/mimi_transformer.py |
| DummyQuantizer | Thin wrapper providing output projection (no true quantization in the inference-only version) that maps encoder output to the exact dimensionality expected by the Flow-LM. | DummyQuantizer in pocket_tts/modules/dummy_quantizer.py |
SEANet Encoder and Decoder
The SEANetEncoder (defined at line 44 of seanet.py) processes raw waveforms using convolutional layers to produce frame-level representations. The SEANetDecoder (line 16) performs the inverse operation, reconstructing waveforms from processed latent vectors. These components handle the high-level frame rate conversion while maintaining temporal fidelity.
ProjectedTransformer
The ProjectedTransformer (line 104 of mimi_transformer.py) acts as a bridge between the SEANet components and the Flow-LM. It projects embeddings into the latent space required for generative modeling, enabling the codec to interface with the language model during both encoding and decoding phases.
DummyQuantizer
The DummyQuantizer (line 5 of dummy_quantizer.py) provides dimensionality matching without performing actual quantization. This design choice reflects the inference-only nature of the current implementation, where the codec operates as a continuous compression system rather than a quantized discrete representation.
Encoding and Decoding API
The MimiModel class exposes two primary methods for audio compression and reconstruction:
-
encode_to_latent(x: torch.Tensor) -> torch.Tensor: Accepts a batch of waveforms with shape[B, C, T](batch, channels, samples), pads them to an integer multiple of the frame size, and processes them through the encoder, encoder-side transformer, and downsampling layers. The output is an unquantized latent tensor ready for the Flow-LM. -
decode_from_latent(latent: torch.Tensor, mimi_state) -> torch.Tensor: Receives a latent tensor produced by the Flow-LM, upsamples it to the encoder frame rate, applies the decoder-side transformer, and passes it through the SEANet decoder to obtain a reconstructed waveform.
Both methods are wrapped with timing utilities (display_execution_time) to monitor performance characteristics during execution.
Integration in the TTS Pipeline
The Mimi codec integrates into the Pocket-TTS pipeline at two critical stages:
-
Audio Prompt Encoding: When processing voice cloning prompts,
TTSModel.get_state_for_audio_promptcalls_encode_audio, which internally invokesself.mimi.encode_to_latent. The resulting latent tensor is linearly projected into the Flow-LM's latent space usingspeaker_proj_weight. -
Latent Decoding: During generation, the Flow-LM produces latent vectors that are fed to
MimiModel.decode_from_latent. This restores the waveform chunk-by-chunk, enabling streaming playback without waiting for complete file generation.
This architecture allows the system to operate efficiently on CPU-only environments while maintaining real-time streaming capabilities.
Practical Code Examples
Encoding and Decoding Audio
The following example demonstrates standalone encoding and decoding using the Mimi codec:
import torch
from pocket_tts import TTSModel
# Load a pre-trained model (weights are downloaded automatically)
model = TTSModel.load_model()
# Encode a raw audio tensor to Mimi latent space
# `audio` must be a FloatTensor of shape [channels, samples]
# Example: a 1-second clip at 24 kHz -> shape [1, 24000]
audio = torch.randn(1, 24000) # replace with real audio data
latent = model.mimi.encode_to_latent(audio.unsqueeze(0)) # batch dim added
print("Latent shape:", latent.shape) # e.g. [1, 80, 250] (B, latent_dim, frames)
# Decode the latent back into a waveform
# Initialise a fresh Mimi state for decoding
mimi_state = model.mimi.init_states(batch_size=1, sequence_length=latent.shape[-1])
# Decode
reconstructed = model.mimi.decode_from_latent(latent, mimi_state)
print("Reconstructed audio shape:", reconstructed.shape) # [1, 1, samples]
Full Voice Cloning Workflow
This example shows the complete pipeline from audio prompt to synthesized speech:
# Load an audio prompt (e.g., a short speaker sample)
voice_state = model.get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Generate speech for a given text, streaming chunks
for chunk in model.generate_audio_stream(voice_state, "Hello, world!"):
# `chunk` is a 1-D Tensor of raw samples ready for playback or saving
# (e.g., write to a WAV file with scipy.io.wavfile.write)
pass
Summary
- The Mimi neural audio codec provides neural compression for raw audio waveforms in the Pocket-TTS system.
- It consists of SEANet convolutional encoder/decoder pairs, a ProjectedTransformer for latent space alignment, and a DummyQuantizer for dimensionality projection.
- The
encode_to_latentmethod inpocket_tts/models/mimi.pycompresses waveforms for the Flow-LM, whiledecode_from_latentreconstructs audio for streaming output. - The codec enables voice cloning by encoding audio prompts and supports real-time generation through chunk-based decoding.
Frequently Asked Questions
How does the Mimi neural audio codec differ from traditional audio codecs?
Traditional audio codecs rely on hand-designed algorithms like psychoacoustic models, whereas the Mimi neural audio codec uses learned convolutional representations through the SEANet architecture. This neural approach allows the system to optimize compression specifically for the downstream Flow-LM generative model rather than human perceptual limits alone.
What is the purpose of the DummyQuantizer in the codec architecture?
The DummyQuantizer serves as a dimensionality adapter that maps the encoder's output to the exact latent dimensionality expected by the Flow-LM. According to the source code in pocket_tts/modules/dummy_quantizer.py, it performs no true quantization in the inference-only version, operating instead as a learned projection layer that maintains continuous representations.
What audio format does the Mimi codec expect for encoding?
The encode_to_latent method expects a FloatTensor of shape [B, C, T] representing batch, channels, and time samples. The code examples demonstrate 24 kHz sampling rates, though the architecture handles various lengths by padding inputs to integer multiples of the frame size before processing through the SEANet encoder.
Is the Mimi codec suitable for real-time streaming applications?
Yes, the codec supports streaming through its decode_from_latent method, which processes audio chunk-by-chunk rather than requiring full sequences. The implementation includes display_execution_time wrappers for performance monitoring, and the architecture is optimized for CPU-only operation, making it suitable for low-latency synthesis scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →