# Complete Guide to Mimi Audio Codec Configuration in PersonaPlex

> Master Mimi audio codec configuration in PersonaPlex. Learn about sample rate, frame rate, and quantizer settings for optimal pipeline assembly. Get the full guide.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: how-to-guide
- Published: 2026-04-07

---

**PersonaPlex configures the Mimi audio codec with a 24 kHz sample rate, 12.5 fps frame rate, and a SplitResidualVectorQuantizer with 32 codebooks, assembling the full pipeline in `get_mimi()` from [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py).**

The **Mimi audio codec** serves as the raw-waveform audio backbone for NVIDIA's PersonaPlex, handling high-fidelity neural audio compression through a sophisticated combination of SEANet encoders, transformer projections, and residual vector quantization. Understanding the exact configuration parameters is essential for anyone modifying the inference pipeline or integrating the codec into custom workflows. All architectural constants and factory methods are defined in the `moshi` submodule, specifically within the model loaders and compression modules.

## Core Configuration Parameters

The **Mimi audio codec configuration** in PersonaPlex relies on several immutable constants and keyword-argument dictionaries defined at the module level in [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py). These settings control the temporal resolution, latent dimensionality, and quantization behavior.

### Sample Rate and Frame Rate

The codec processes mono audio at **24 kHz** (`SAMPLE_RATE = 24000`) and produces discrete tokens at **12.5 frames per second** (`FRAME_RATE = 12.5`), corresponding to approximately 192 ms per frame. These constants are declared on lines 39-42 of [`loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/loaders.py) and directly influence the hop length calculations in the SEANet encoder.

### SEANet Encoder and Decoder Architecture

Both the encoder and decoder share identical hyper-parameters via the `_seanet_kwargs` dictionary (lines 48-67). The configuration specifies:

- **Single-channel** input/output audio
- **512-dimensional** latent space
- **Causal** convolutional processing
- **64 filters** with **dilation = 2**
- Strided convolutions that determine the encoder frame rate (`SAMPLE_RATE / hop_length`)

### Transformer Projections

The encoder and decoder each utilize a `ProjectedTransformer` configured via `_transformer_kwargs` (lines 34-44). Key specifications include:

- **512-dimensional** model size
- **8 attention heads** across **8 layers**
- **Causal** masking with RoPE (Rotary Positional Embeddings)
- **Layer-scale** initialization of 0.01

### Vector Quantization Setup

Quantization employs a `SplitResidualVectorQuantizer` defined in `_quantizer_kwargs` (lines 68-74):

- **256-dimensional** embeddings
- **32 total codebooks** with **2048 bins** each
- Input and output projections of **256 dimensions**

## Model Assembly in `get_mimi()`

The `get_mimi()` factory function (lines 29-55 in [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py)) assembles these components into a functional `MimiModel` instance. This function:

1. Instantiates `SEANetEncoder` and `SEANetDecoder` with `_seanet_kwargs`
2. Builds projected transformers for both encoder and decoder pathways
3. Initializes the `SplitResidualVectorQuantizer` with `_quantizer_kwargs`
4. Assembles the full pipeline via `MimiModel` (defined in [`moshi/moshi/models/compression.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/compression.py)) with `causal=True` and `resample_method="conv"`
5. Loads the **Safetensors or PyTorch checkpoint** specified by the `MIMI_NAME` constant (`tokenizer-e351c8d8-checkpoint125.safetensors` on line 44)
6. Activates only **8 codebooks** via `model.set_num_codebooks(8)` for inference

The resulting model operates strictly in evaluation mode (`model.eval()`) with device placement handled via the `device` parameter.

## Practical Code Examples

### Loading the Codec

```python
from moshi.moshi.models.loaders import get_mimi
from pathlib import Path

ckpt_path = Path("tokenizer-e351c8d8-checkpoint125.safetensors")
mimi = get_mimi(ckpt_path, device="cpu")

```

### Encoding Audio to Discrete Tokens

```python
import torch

# Generate 1 second of mono audio at 24 kHz

waveform = torch.randn(1, 1, 24000)
tokens = mimi.encode(waveform)

# Returns shape [batch, active_codebooks, time_frames]

print(tokens.shape)  # torch.Size([1, 8, 300])

```

### Decoding Tokens to Waveform

```python
reconstructed = mimi.decode(tokens)
print(reconstructed.shape)  # torch.Size([1, 1, 24000])

```

### Accessing Unquantized Latents

```python
latent = mimi._encode_to_unquantized_latent(waveform)

# Shape: [batch, 512, time_frames] float tensor

reconstructed = mimi.decode_latent(mimi.quantizer.encode(latent))

```

## Summary

- **Configuration Location**: All constants reside in [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py), while the `MimiModel` class implementation lives in [`moshi/moshi/models/compression.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/compression.py).
- **Temporal Settings**: Fixed at 24 kHz sample rate with 12.5 fps token rate, yielding ~192 ms per frame.
- **Architecture**: SEANet encoder/decoder with 512-dim latents, 8-layer transformers with RoPE, and 32-codebook residual quantization (8 active during inference).
- **Checkpoint**: Defaults to `tokenizer-e351c8d8-checkpoint125.safetensors` loaded through the `get_mimi()` factory.
- **Causal Processing**: The entire pipeline enforces causal masking, making it suitable for real-time streaming applications.

## Frequently Asked Questions

### What is the exact sample rate and frame rate for Mimi in PersonaPlex?

The Mimi codec operates at **24 kHz** (`SAMPLE_RATE = 24000`) with a token frame rate of **12.5 fps** (`FRAME_RATE = 12.5`). According to the source code in [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py), this results in approximately 192 milliseconds of audio per discrete token frame.

### How many codebooks does the Mimi quantizer use, and how many are active?

The `SplitResidualVectorQuantizer` is configured with **32 total codebooks**, each containing **2048 bins**. However, during model initialization in `get_mimi()`, only **8 codebooks** are activated via `model.set_num_codebooks(8)` for standard inference, reducing bandwidth while maintaining quality.

### Where is the Mimi model assembled in the PersonaPlex codebase?

The complete model assembly occurs in the **`get_mimi()` function** inside [`moshi/moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/loaders.py) (lines 29-55). This factory instantiates the SEANet encoder/decoder, projected transformers, and quantizer before wrapping them in the `MimiModel` class imported from [`moshi/moshi/models/compression.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/compression.py).

### What checkpoint file is required to initialize the Mimi codec?

The default checkpoint is defined by the `MIMI_NAME` constant as **`tokenizer-e351c8d8-checkpoint125.safetensors`** (line 44 of [`loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/loaders.py)). The `get_mimi()` function automatically loads this Safetensors format (or legacy PyTorch checkpoints) and applies the weights to the assembled model.