Complete Guide to Mimi Audio Codec Configuration in PersonaPlex
PersonaPlex configures the Mimi audio codec with a 24 kHz sample rate, 12.5 fps frame rate, and a SplitResidualVectorQuantizer with 32 codebooks, assembling the full pipeline in get_mimi() from moshi/moshi/models/loaders.py.
The Mimi audio codec serves as the raw-waveform audio backbone for NVIDIA's PersonaPlex, handling high-fidelity neural audio compression through a sophisticated combination of SEANet encoders, transformer projections, and residual vector quantization. Understanding the exact configuration parameters is essential for anyone modifying the inference pipeline or integrating the codec into custom workflows. All architectural constants and factory methods are defined in the moshi submodule, specifically within the model loaders and compression modules.
Core Configuration Parameters
The Mimi audio codec configuration in PersonaPlex relies on several immutable constants and keyword-argument dictionaries defined at the module level in moshi/moshi/models/loaders.py. These settings control the temporal resolution, latent dimensionality, and quantization behavior.
Sample Rate and Frame Rate
The codec processes mono audio at 24 kHz (SAMPLE_RATE = 24000) and produces discrete tokens at 12.5 frames per second (FRAME_RATE = 12.5), corresponding to approximately 192 ms per frame. These constants are declared on lines 39-42 of loaders.py and directly influence the hop length calculations in the SEANet encoder.
SEANet Encoder and Decoder Architecture
Both the encoder and decoder share identical hyper-parameters via the _seanet_kwargs dictionary (lines 48-67). The configuration specifies:
- Single-channel input/output audio
- 512-dimensional latent space
- Causal convolutional processing
- 64 filters with dilation = 2
- Strided convolutions that determine the encoder frame rate (
SAMPLE_RATE / hop_length)
Transformer Projections
The encoder and decoder each utilize a ProjectedTransformer configured via _transformer_kwargs (lines 34-44). Key specifications include:
- 512-dimensional model size
- 8 attention heads across 8 layers
- Causal masking with RoPE (Rotary Positional Embeddings)
- Layer-scale initialization of 0.01
Vector Quantization Setup
Quantization employs a SplitResidualVectorQuantizer defined in _quantizer_kwargs (lines 68-74):
- 256-dimensional embeddings
- 32 total codebooks with 2048 bins each
- Input and output projections of 256 dimensions
Model Assembly in get_mimi()
The get_mimi() factory function (lines 29-55 in moshi/moshi/models/loaders.py) assembles these components into a functional MimiModel instance. This function:
- Instantiates
SEANetEncoderandSEANetDecoderwith_seanet_kwargs - Builds projected transformers for both encoder and decoder pathways
- Initializes the
SplitResidualVectorQuantizerwith_quantizer_kwargs - Assembles the full pipeline via
MimiModel(defined inmoshi/moshi/models/compression.py) withcausal=Trueandresample_method="conv" - Loads the Safetensors or PyTorch checkpoint specified by the
MIMI_NAMEconstant (tokenizer-e351c8d8-checkpoint125.safetensorson line 44) - Activates only 8 codebooks via
model.set_num_codebooks(8)for inference
The resulting model operates strictly in evaluation mode (model.eval()) with device placement handled via the device parameter.
Practical Code Examples
Loading the Codec
from moshi.moshi.models.loaders import get_mimi
from pathlib import Path
ckpt_path = Path("tokenizer-e351c8d8-checkpoint125.safetensors")
mimi = get_mimi(ckpt_path, device="cpu")
Encoding Audio to Discrete Tokens
import torch
# Generate 1 second of mono audio at 24 kHz
waveform = torch.randn(1, 1, 24000)
tokens = mimi.encode(waveform)
# Returns shape [batch, active_codebooks, time_frames]
print(tokens.shape) # torch.Size([1, 8, 300])
Decoding Tokens to Waveform
reconstructed = mimi.decode(tokens)
print(reconstructed.shape) # torch.Size([1, 1, 24000])
Accessing Unquantized Latents
latent = mimi._encode_to_unquantized_latent(waveform)
# Shape: [batch, 512, time_frames] float tensor
reconstructed = mimi.decode_latent(mimi.quantizer.encode(latent))
Summary
- Configuration Location: All constants reside in
moshi/moshi/models/loaders.py, while theMimiModelclass implementation lives inmoshi/moshi/models/compression.py. - Temporal Settings: Fixed at 24 kHz sample rate with 12.5 fps token rate, yielding ~192 ms per frame.
- Architecture: SEANet encoder/decoder with 512-dim latents, 8-layer transformers with RoPE, and 32-codebook residual quantization (8 active during inference).
- Checkpoint: Defaults to
tokenizer-e351c8d8-checkpoint125.safetensorsloaded through theget_mimi()factory. - Causal Processing: The entire pipeline enforces causal masking, making it suitable for real-time streaming applications.
Frequently Asked Questions
What is the exact sample rate and frame rate for Mimi in PersonaPlex?
The Mimi codec operates at 24 kHz (SAMPLE_RATE = 24000) with a token frame rate of 12.5 fps (FRAME_RATE = 12.5). According to the source code in moshi/moshi/models/loaders.py, this results in approximately 192 milliseconds of audio per discrete token frame.
How many codebooks does the Mimi quantizer use, and how many are active?
The SplitResidualVectorQuantizer is configured with 32 total codebooks, each containing 2048 bins. However, during model initialization in get_mimi(), only 8 codebooks are activated via model.set_num_codebooks(8) for standard inference, reducing bandwidth while maintaining quality.
Where is the Mimi model assembled in the PersonaPlex codebase?
The complete model assembly occurs in the get_mimi() function inside moshi/moshi/models/loaders.py (lines 29-55). This factory instantiates the SEANet encoder/decoder, projected transformers, and quantizer before wrapping them in the MimiModel class imported from moshi/moshi/models/compression.py.
What checkpoint file is required to initialize the Mimi codec?
The default checkpoint is defined by the MIMI_NAME constant as tokenizer-e351c8d8-checkpoint125.safetensors (line 44 of loaders.py). The get_mimi() function automatically loads this Safetensors format (or legacy PyTorch checkpoints) and applies the weights to the assembled model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →