# Understanding Audio Frame Rate and Token Structure in NVIDIA PersonaPlex

> Explore NVIDIA PersonaPlex audio processing at 12.5 Hz. Understand token structure with 8 audio and 1 text token for mixed-modality streams managed by LM and Mimi codec.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: deep-dive
- Published: 2026-04-07

---

**PersonaPlex processes audio at 12.5 Hz (80 ms frames) with a 24 kHz sample rate, encoding each frame into 8 parallel audio tokens plus one text token to create a mixed-modality stream managed by the LM and Mimi codec.**

The **audio frame rate and token structure in PersonaPlex** define how this open-source multimodal system synchronizes text and audio generation. PersonaPlex uses a fixed-frame approach where the language model (LM) predicts discrete tokens representing both linguistic content and compressed audio waveforms. Understanding these constants—defined across [`moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/loaders.py), [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py), and the server implementations—is essential for customizing inference or debugging token streams.

## Audio Frame Rate Constants and Calculation

PersonaPlex operates on fixed-size audio frames derived from two core sampling constants that bridge raw audio and tokenized representations.

### Core Sampling Parameters

The system defines temporal granularity through constants declared in [`moshi/models/loaders.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/loaders.py):

- **`SAMPLE_RATE`**: Set to `24000` (24 kHz) at line 39, representing the raw audio sampling frequency
- **`FRAME_RATE`**: Set to `12.5` Hz (12.5 frames per second) at line 40, corresponding to 80 ms per frame
- **`FRAME_RATE_HZ`**: A duplicate constant (`12.5`) defined at `moshi/models/lm.py#L55` for use within the language model

### Frame Size Computation

The **frame size in samples** is computed dynamically by dividing the sample rate by the frame rate:

```python
self.frame_size = int(self.mimi.sample_rate / self.mimi.frame_rate)

# Result: 24000 / 12.5 = 1920 samples (0.08 seconds)

```

This calculation appears in the server implementation at `moshi/server.py#L105-L107` and the offline runner at `moshi/offline.py#L211-L214`. The 12.5 Hz frame rate matches the temporal granularity used by the **Mimi** audio codec, ensuring each frame corresponds to a chunk of audio that the model encodes and decodes as discrete tokens.

## Mixed Modality Token Architecture

PersonaPlex employs a **mixed modality token stream** combining text symbols with parallel audio codebooks at each generation step.

### Parallel Audio Streams

The token structure relies on several dimensionality constants defined in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py):

- **`dep_q`**: The number of independent audio codebooks (parallel streams), defaulting to `8` at line 246
- **`AUDIO_TOKENS_PER_STREAM`**: Fixed at `8` tokens per frame per stream at line 54
- **Total token dimension**: `dep_q + 1` (8 audio streams + 1 text token), validated at `moshi/server.py#L230` via `assert tokens.shape[1] == self.lm_gen.lm_model.dep_q + 1`

At each generation step, the model produces a tensor of shape `[B, dep_q + 1, 1]`, where the first slot contains the text token and the remaining slots contain the audio codebook tokens for the current frame.

### Special Token Sets

PersonaPlex defines specific token IDs for special audio conditions and text boundaries:

**Silence tokens** encode a short silent audio chunk:

```python
SILENCE_TOKENS = np.array([948, 243, 1178, 546, 1736, 1030, 1978, 2008], dtype=np.int64)

```

Defined at `moshi/models/lm.py#L56`.

**Sine-wave tokens** synthesize a pure tone for testing or filler audio:

```python
SINE_TOKENS = np.array([430, 1268, 381, 1611, 1095, 1495, 56, 472], dtype=np.int64)

```

Defined at `moshi/models/lm.py#L57`.

**Text special tokens** are mapped to human-readable control symbols:

```python
text_token_map = ['EPAD', 'BOS', 'EOS', 'PAD']

```

These mappings are used during server decoding at `moshi/server.py#L242-L243` and offline processing at `moshi/offline.py#L293-L294` to identify control tokens versus content tokens.

### Zero-Token Padding

The model reserves a **zero-token** (exposed via `zero_token_id()` at `moshi/models/lm.py#L379-L380`) to indicate "no sampling" or padding positions. This token masks unused positions in the attention matrix and during loss computation (lines 549-551 in [`lm.py`](https://github.com/NVIDIA/personaplex/blob/main/lm.py)).

## Token Generation and Decoding Flow

The inference pipeline moves from text prompting to audio reconstruction through a structured token exchange:

1. **Initialization**: The LM generates forced tokens containing text prompts and optional audio priming via `lm_gen.step`
2. **Sampling**: The `sample_token` function (in [`utils/sampling.py`](https://github.com/NVIDIA/personaplex/blob/main/utils/sampling.py)) draws the next token for each stream using top-k, top-p, or greedy selection
3. **Packaging**: Tokens assemble into a tensor of shape `[B, dep_q + 1, 1]` where index 0 is text and indices 1-8 are audio
4. **Decoding**: The `MimiModel` ([`models/compression.py`](https://github.com/NVIDIA/personaplex/blob/main/models/compression.py)) receives audio tokens, expands them back to PCM, and writes the result at the appropriate sample offset using the 1920-sample frame size

### Practical Usage Example

```python

# Assume lm_gen is an LMGen instance and mimi is a MimiModel

text_prompt = "Hello, world!"
lm_gen.text_prompt_tokens = tokenizer.encode(wrap_with_system_tags(text_prompt))

while not done:
    # Generate next token block: shape [B, dep_q+1, 1]

    tokens = lm_gen.step(prev_codes)
    
    # Extract text token for logging

    text_token_id = tokens[0, 0, 0].item()
    if text_token_id not in (0, 3):  # Ignore PAD/EOS

        print("Text:", tokenizer.id_to_piece(text_token_id))
    
    # Extract audio tokens: shape [B, dep_q, 1]

    audio_tokens = tokens[:, 1:, :]
    
    # Decode to PCM frame (1920 samples at 24kHz)

    pcm_frame = mimi.decode(audio_tokens[:, :, 0])
    output_pcm.append(pcm_frame)
    
    prev_codes = tokens

```

## Summary

- **Audio frame rate** is fixed at **12.5 Hz** (80 ms frames) with a **24 kHz sample rate**, yielding 1920 samples per frame as calculated in [`server.py`](https://github.com/NVIDIA/personaplex/blob/main/server.py) and [`offline.py`](https://github.com/NVIDIA/personaplex/blob/main/offline.py)
- **Token structure** consists of **8 parallel audio codebooks** (`dep_q=8`) plus **1 text token** per step, totaling 9 tokens per frame
- **Special tokens** include silence arrays (`SILENCE_TOKENS`), sine-wave arrays (`SINE_TOKENS`), and text control symbols (BOS/EOS/PAD/EPAD) defined in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py)
- **Zero-token** handling masks padding positions during attention and loss computation via `zero_token_id()`
- The **Mimi codec** reconstructs PCM audio from the 8 parallel audio token streams at the specified frame boundaries

## Frequently Asked Questions

### What is the audio frame duration in PersonaPlex?

Each audio frame represents **80 milliseconds** of audio. This duration derives from the `FRAME_RATE` of 12.5 Hz defined in `moshi/models/loaders.py#L40`, calculated as 1/12.5 = 0.08 seconds. With a sample rate of 24 kHz, each frame contains exactly 1920 samples.

### How many audio tokens does PersonaPlex generate per frame?

PersonaPlex generates **8 audio tokens per frame**, one for each parallel codebook stream (`dep_q=8`). These tokens are produced alongside a single text token, resulting in a total of 9 tokens per generation step. The audio tokens are fed to the Mimi codec for reconstruction into PCM waveforms.

### What are silence tokens used for in PersonaPlex?

**Silence tokens** are pre-defined token IDs (`[948, 243, 1178, 546, 1736, 1030, 1978, 2008]`) that encode a silent audio chunk when decoded by the Mimi model. Defined in `moshi/models/lm.py#L56`, these tokens allow the model to explicitly represent pauses or non-speech intervals in the audio stream without predicting raw audio samples.

### Where is the zero-token defined and what does it do?

The **zero-token** is defined in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py) via the `zero_token_id()` method (lines 379-380). This token reserves ID 0 to represent padding or "no sampling" positions in the attention matrix. It masks out unused positions during training loss computation (lines 549-551) and prevents the model from attending to or predicting content at placeholder locations in the sequence.