Complete Guide to Mimi Audio Codec Configuration in PersonaPlex

PersonaPlex configures the Mimi audio codec with a 24 kHz sample rate, 12.5 fps frame rate, and a SplitResidualVectorQuantizer with 32 codebooks, assembling the full pipeline in get_mimi() from moshi/moshi/models/loaders.py.

The Mimi audio codec serves as the raw-waveform audio backbone for NVIDIA's PersonaPlex, handling high-fidelity neural audio compression through a sophisticated combination of SEANet encoders, transformer projections, and residual vector quantization. Understanding the exact configuration parameters is essential for anyone modifying the inference pipeline or integrating the codec into custom workflows. All architectural constants and factory methods are defined in the moshi submodule, specifically within the model loaders and compression modules.

Core Configuration Parameters

The Mimi audio codec configuration in PersonaPlex relies on several immutable constants and keyword-argument dictionaries defined at the module level in moshi/moshi/models/loaders.py. These settings control the temporal resolution, latent dimensionality, and quantization behavior.

Sample Rate and Frame Rate

The codec processes mono audio at 24 kHz (SAMPLE_RATE = 24000) and produces discrete tokens at 12.5 frames per second (FRAME_RATE = 12.5), corresponding to approximately 192 ms per frame. These constants are declared on lines 39-42 of loaders.py and directly influence the hop length calculations in the SEANet encoder.

SEANet Encoder and Decoder Architecture

Both the encoder and decoder share identical hyper-parameters via the _seanet_kwargs dictionary (lines 48-67). The configuration specifies:

  • Single-channel input/output audio
  • 512-dimensional latent space
  • Causal convolutional processing
  • 64 filters with dilation = 2
  • Strided convolutions that determine the encoder frame rate (SAMPLE_RATE / hop_length)

Transformer Projections

The encoder and decoder each utilize a ProjectedTransformer configured via _transformer_kwargs (lines 34-44). Key specifications include:

  • 512-dimensional model size
  • 8 attention heads across 8 layers
  • Causal masking with RoPE (Rotary Positional Embeddings)
  • Layer-scale initialization of 0.01

Vector Quantization Setup

Quantization employs a SplitResidualVectorQuantizer defined in _quantizer_kwargs (lines 68-74):

  • 256-dimensional embeddings
  • 32 total codebooks with 2048 bins each
  • Input and output projections of 256 dimensions

Model Assembly in get_mimi()

The get_mimi() factory function (lines 29-55 in moshi/moshi/models/loaders.py) assembles these components into a functional MimiModel instance. This function:

  1. Instantiates SEANetEncoder and SEANetDecoder with _seanet_kwargs
  2. Builds projected transformers for both encoder and decoder pathways
  3. Initializes the SplitResidualVectorQuantizer with _quantizer_kwargs
  4. Assembles the full pipeline via MimiModel (defined in moshi/moshi/models/compression.py) with causal=True and resample_method="conv"
  5. Loads the Safetensors or PyTorch checkpoint specified by the MIMI_NAME constant (tokenizer-e351c8d8-checkpoint125.safetensors on line 44)
  6. Activates only 8 codebooks via model.set_num_codebooks(8) for inference

The resulting model operates strictly in evaluation mode (model.eval()) with device placement handled via the device parameter.

Practical Code Examples

Loading the Codec

from moshi.moshi.models.loaders import get_mimi
from pathlib import Path

ckpt_path = Path("tokenizer-e351c8d8-checkpoint125.safetensors")
mimi = get_mimi(ckpt_path, device="cpu")

Encoding Audio to Discrete Tokens

import torch

# Generate 1 second of mono audio at 24 kHz

waveform = torch.randn(1, 1, 24000)
tokens = mimi.encode(waveform)

# Returns shape [batch, active_codebooks, time_frames]

print(tokens.shape)  # torch.Size([1, 8, 300])

Decoding Tokens to Waveform

reconstructed = mimi.decode(tokens)
print(reconstructed.shape)  # torch.Size([1, 1, 24000])

Accessing Unquantized Latents

latent = mimi._encode_to_unquantized_latent(waveform)

# Shape: [batch, 512, time_frames] float tensor

reconstructed = mimi.decode_latent(mimi.quantizer.encode(latent))

Summary

  • Configuration Location: All constants reside in moshi/moshi/models/loaders.py, while the MimiModel class implementation lives in moshi/moshi/models/compression.py.
  • Temporal Settings: Fixed at 24 kHz sample rate with 12.5 fps token rate, yielding ~192 ms per frame.
  • Architecture: SEANet encoder/decoder with 512-dim latents, 8-layer transformers with RoPE, and 32-codebook residual quantization (8 active during inference).
  • Checkpoint: Defaults to tokenizer-e351c8d8-checkpoint125.safetensors loaded through the get_mimi() factory.
  • Causal Processing: The entire pipeline enforces causal masking, making it suitable for real-time streaming applications.

Frequently Asked Questions

What is the exact sample rate and frame rate for Mimi in PersonaPlex?

The Mimi codec operates at 24 kHz (SAMPLE_RATE = 24000) with a token frame rate of 12.5 fps (FRAME_RATE = 12.5). According to the source code in moshi/moshi/models/loaders.py, this results in approximately 192 milliseconds of audio per discrete token frame.

How many codebooks does the Mimi quantizer use, and how many are active?

The SplitResidualVectorQuantizer is configured with 32 total codebooks, each containing 2048 bins. However, during model initialization in get_mimi(), only 8 codebooks are activated via model.set_num_codebooks(8) for standard inference, reducing bandwidth while maintaining quality.

Where is the Mimi model assembled in the PersonaPlex codebase?

The complete model assembly occurs in the get_mimi() function inside moshi/moshi/models/loaders.py (lines 29-55). This factory instantiates the SEANet encoder/decoder, projected transformers, and quantizer before wrapping them in the MimiModel class imported from moshi/moshi/models/compression.py.

What checkpoint file is required to initialize the Mimi codec?

The default checkpoint is defined by the MIMI_NAME constant as tokenizer-e351c8d8-checkpoint125.safetensors (line 44 of loaders.py). The get_mimi() function automatically loads this Safetensors format (or legacy PyTorch checkpoints) and applies the weights to the assembled model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →