Understanding Audio Frame Rate and Token Structure in NVIDIA PersonaPlex

PersonaPlex processes audio at 12.5 Hz (80 ms frames) with a 24 kHz sample rate, encoding each frame into 8 parallel audio tokens plus one text token to create a mixed-modality stream managed by the LM and Mimi codec.

The audio frame rate and token structure in PersonaPlex define how this open-source multimodal system synchronizes text and audio generation. PersonaPlex uses a fixed-frame approach where the language model (LM) predicts discrete tokens representing both linguistic content and compressed audio waveforms. Understanding these constants—defined across moshi/models/loaders.py, moshi/models/lm.py, and the server implementations—is essential for customizing inference or debugging token streams.

Audio Frame Rate Constants and Calculation

PersonaPlex operates on fixed-size audio frames derived from two core sampling constants that bridge raw audio and tokenized representations.

Core Sampling Parameters

The system defines temporal granularity through constants declared in moshi/models/loaders.py:

  • SAMPLE_RATE: Set to 24000 (24 kHz) at line 39, representing the raw audio sampling frequency
  • FRAME_RATE: Set to 12.5 Hz (12.5 frames per second) at line 40, corresponding to 80 ms per frame
  • FRAME_RATE_HZ: A duplicate constant (12.5) defined at moshi/models/lm.py#L55 for use within the language model

Frame Size Computation

The frame size in samples is computed dynamically by dividing the sample rate by the frame rate:

self.frame_size = int(self.mimi.sample_rate / self.mimi.frame_rate)

# Result: 24000 / 12.5 = 1920 samples (0.08 seconds)

This calculation appears in the server implementation at moshi/server.py#L105-L107 and the offline runner at moshi/offline.py#L211-L214. The 12.5 Hz frame rate matches the temporal granularity used by the Mimi audio codec, ensuring each frame corresponds to a chunk of audio that the model encodes and decodes as discrete tokens.

Mixed Modality Token Architecture

PersonaPlex employs a mixed modality token stream combining text symbols with parallel audio codebooks at each generation step.

Parallel Audio Streams

The token structure relies on several dimensionality constants defined in moshi/models/lm.py:

  • dep_q: The number of independent audio codebooks (parallel streams), defaulting to 8 at line 246
  • AUDIO_TOKENS_PER_STREAM: Fixed at 8 tokens per frame per stream at line 54
  • Total token dimension: dep_q + 1 (8 audio streams + 1 text token), validated at moshi/server.py#L230 via assert tokens.shape[1] == self.lm_gen.lm_model.dep_q + 1

At each generation step, the model produces a tensor of shape [B, dep_q + 1, 1], where the first slot contains the text token and the remaining slots contain the audio codebook tokens for the current frame.

Special Token Sets

PersonaPlex defines specific token IDs for special audio conditions and text boundaries:

Silence tokens encode a short silent audio chunk:

SILENCE_TOKENS = np.array([948, 243, 1178, 546, 1736, 1030, 1978, 2008], dtype=np.int64)

Defined at moshi/models/lm.py#L56.

Sine-wave tokens synthesize a pure tone for testing or filler audio:

SINE_TOKENS = np.array([430, 1268, 381, 1611, 1095, 1495, 56, 472], dtype=np.int64)

Defined at moshi/models/lm.py#L57.

Text special tokens are mapped to human-readable control symbols:

text_token_map = ['EPAD', 'BOS', 'EOS', 'PAD']

These mappings are used during server decoding at moshi/server.py#L242-L243 and offline processing at moshi/offline.py#L293-L294 to identify control tokens versus content tokens.

Zero-Token Padding

The model reserves a zero-token (exposed via zero_token_id() at moshi/models/lm.py#L379-L380) to indicate "no sampling" or padding positions. This token masks unused positions in the attention matrix and during loss computation (lines 549-551 in lm.py).

Token Generation and Decoding Flow

The inference pipeline moves from text prompting to audio reconstruction through a structured token exchange:

  1. Initialization: The LM generates forced tokens containing text prompts and optional audio priming via lm_gen.step
  2. Sampling: The sample_token function (in utils/sampling.py) draws the next token for each stream using top-k, top-p, or greedy selection
  3. Packaging: Tokens assemble into a tensor of shape [B, dep_q + 1, 1] where index 0 is text and indices 1-8 are audio
  4. Decoding: The MimiModel (models/compression.py) receives audio tokens, expands them back to PCM, and writes the result at the appropriate sample offset using the 1920-sample frame size

Practical Usage Example


# Assume lm_gen is an LMGen instance and mimi is a MimiModel

text_prompt = "Hello, world!"
lm_gen.text_prompt_tokens = tokenizer.encode(wrap_with_system_tags(text_prompt))

while not done:
    # Generate next token block: shape [B, dep_q+1, 1]

    tokens = lm_gen.step(prev_codes)
    
    # Extract text token for logging

    text_token_id = tokens[0, 0, 0].item()
    if text_token_id not in (0, 3):  # Ignore PAD/EOS

        print("Text:", tokenizer.id_to_piece(text_token_id))
    
    # Extract audio tokens: shape [B, dep_q, 1]

    audio_tokens = tokens[:, 1:, :]
    
    # Decode to PCM frame (1920 samples at 24kHz)

    pcm_frame = mimi.decode(audio_tokens[:, :, 0])
    output_pcm.append(pcm_frame)
    
    prev_codes = tokens

Summary

  • Audio frame rate is fixed at 12.5 Hz (80 ms frames) with a 24 kHz sample rate, yielding 1920 samples per frame as calculated in server.py and offline.py
  • Token structure consists of 8 parallel audio codebooks (dep_q=8) plus 1 text token per step, totaling 9 tokens per frame
  • Special tokens include silence arrays (SILENCE_TOKENS), sine-wave arrays (SINE_TOKENS), and text control symbols (BOS/EOS/PAD/EPAD) defined in moshi/models/lm.py
  • Zero-token handling masks padding positions during attention and loss computation via zero_token_id()
  • The Mimi codec reconstructs PCM audio from the 8 parallel audio token streams at the specified frame boundaries

Frequently Asked Questions

What is the audio frame duration in PersonaPlex?

Each audio frame represents 80 milliseconds of audio. This duration derives from the FRAME_RATE of 12.5 Hz defined in moshi/models/loaders.py#L40, calculated as 1/12.5 = 0.08 seconds. With a sample rate of 24 kHz, each frame contains exactly 1920 samples.

How many audio tokens does PersonaPlex generate per frame?

PersonaPlex generates 8 audio tokens per frame, one for each parallel codebook stream (dep_q=8). These tokens are produced alongside a single text token, resulting in a total of 9 tokens per generation step. The audio tokens are fed to the Mimi codec for reconstruction into PCM waveforms.

What are silence tokens used for in PersonaPlex?

Silence tokens are pre-defined token IDs ([948, 243, 1178, 546, 1736, 1030, 1978, 2008]) that encode a silent audio chunk when decoded by the Mimi model. Defined in moshi/models/lm.py#L56, these tokens allow the model to explicitly represent pauses or non-speech intervals in the audio stream without predicting raw audio samples.

Where is the zero-token defined and what does it do?

The zero-token is defined in moshi/models/lm.py via the zero_token_id() method (lines 379-380). This token reserves ID 0 to represent padding or "no sampling" positions in the attention matrix. It masks out unused positions during training loss computation (lines 549-551) and prevents the model from attending to or predicting content at placeholder locations in the sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →