Understanding Audio Frame Rate and Token Structure in NVIDIA PersonaPlex
PersonaPlex processes audio at 12.5 Hz (80 ms frames) with a 24 kHz sample rate, encoding each frame into 8 parallel audio tokens plus one text token to create a mixed-modality stream managed by the LM and Mimi codec.
The audio frame rate and token structure in PersonaPlex define how this open-source multimodal system synchronizes text and audio generation. PersonaPlex uses a fixed-frame approach where the language model (LM) predicts discrete tokens representing both linguistic content and compressed audio waveforms. Understanding these constants—defined across moshi/models/loaders.py, moshi/models/lm.py, and the server implementations—is essential for customizing inference or debugging token streams.
Audio Frame Rate Constants and Calculation
PersonaPlex operates on fixed-size audio frames derived from two core sampling constants that bridge raw audio and tokenized representations.
Core Sampling Parameters
The system defines temporal granularity through constants declared in moshi/models/loaders.py:
SAMPLE_RATE: Set to24000(24 kHz) at line 39, representing the raw audio sampling frequencyFRAME_RATE: Set to12.5Hz (12.5 frames per second) at line 40, corresponding to 80 ms per frameFRAME_RATE_HZ: A duplicate constant (12.5) defined atmoshi/models/lm.py#L55for use within the language model
Frame Size Computation
The frame size in samples is computed dynamically by dividing the sample rate by the frame rate:
self.frame_size = int(self.mimi.sample_rate / self.mimi.frame_rate)
# Result: 24000 / 12.5 = 1920 samples (0.08 seconds)
This calculation appears in the server implementation at moshi/server.py#L105-L107 and the offline runner at moshi/offline.py#L211-L214. The 12.5 Hz frame rate matches the temporal granularity used by the Mimi audio codec, ensuring each frame corresponds to a chunk of audio that the model encodes and decodes as discrete tokens.
Mixed Modality Token Architecture
PersonaPlex employs a mixed modality token stream combining text symbols with parallel audio codebooks at each generation step.
Parallel Audio Streams
The token structure relies on several dimensionality constants defined in moshi/models/lm.py:
dep_q: The number of independent audio codebooks (parallel streams), defaulting to8at line 246AUDIO_TOKENS_PER_STREAM: Fixed at8tokens per frame per stream at line 54- Total token dimension:
dep_q + 1(8 audio streams + 1 text token), validated atmoshi/server.py#L230viaassert tokens.shape[1] == self.lm_gen.lm_model.dep_q + 1
At each generation step, the model produces a tensor of shape [B, dep_q + 1, 1], where the first slot contains the text token and the remaining slots contain the audio codebook tokens for the current frame.
Special Token Sets
PersonaPlex defines specific token IDs for special audio conditions and text boundaries:
Silence tokens encode a short silent audio chunk:
SILENCE_TOKENS = np.array([948, 243, 1178, 546, 1736, 1030, 1978, 2008], dtype=np.int64)
Defined at moshi/models/lm.py#L56.
Sine-wave tokens synthesize a pure tone for testing or filler audio:
SINE_TOKENS = np.array([430, 1268, 381, 1611, 1095, 1495, 56, 472], dtype=np.int64)
Defined at moshi/models/lm.py#L57.
Text special tokens are mapped to human-readable control symbols:
text_token_map = ['EPAD', 'BOS', 'EOS', 'PAD']
These mappings are used during server decoding at moshi/server.py#L242-L243 and offline processing at moshi/offline.py#L293-L294 to identify control tokens versus content tokens.
Zero-Token Padding
The model reserves a zero-token (exposed via zero_token_id() at moshi/models/lm.py#L379-L380) to indicate "no sampling" or padding positions. This token masks unused positions in the attention matrix and during loss computation (lines 549-551 in lm.py).
Token Generation and Decoding Flow
The inference pipeline moves from text prompting to audio reconstruction through a structured token exchange:
- Initialization: The LM generates forced tokens containing text prompts and optional audio priming via
lm_gen.step - Sampling: The
sample_tokenfunction (inutils/sampling.py) draws the next token for each stream using top-k, top-p, or greedy selection - Packaging: Tokens assemble into a tensor of shape
[B, dep_q + 1, 1]where index 0 is text and indices 1-8 are audio - Decoding: The
MimiModel(models/compression.py) receives audio tokens, expands them back to PCM, and writes the result at the appropriate sample offset using the 1920-sample frame size
Practical Usage Example
# Assume lm_gen is an LMGen instance and mimi is a MimiModel
text_prompt = "Hello, world!"
lm_gen.text_prompt_tokens = tokenizer.encode(wrap_with_system_tags(text_prompt))
while not done:
# Generate next token block: shape [B, dep_q+1, 1]
tokens = lm_gen.step(prev_codes)
# Extract text token for logging
text_token_id = tokens[0, 0, 0].item()
if text_token_id not in (0, 3): # Ignore PAD/EOS
print("Text:", tokenizer.id_to_piece(text_token_id))
# Extract audio tokens: shape [B, dep_q, 1]
audio_tokens = tokens[:, 1:, :]
# Decode to PCM frame (1920 samples at 24kHz)
pcm_frame = mimi.decode(audio_tokens[:, :, 0])
output_pcm.append(pcm_frame)
prev_codes = tokens
Summary
- Audio frame rate is fixed at 12.5 Hz (80 ms frames) with a 24 kHz sample rate, yielding 1920 samples per frame as calculated in
server.pyandoffline.py - Token structure consists of 8 parallel audio codebooks (
dep_q=8) plus 1 text token per step, totaling 9 tokens per frame - Special tokens include silence arrays (
SILENCE_TOKENS), sine-wave arrays (SINE_TOKENS), and text control symbols (BOS/EOS/PAD/EPAD) defined inmoshi/models/lm.py - Zero-token handling masks padding positions during attention and loss computation via
zero_token_id() - The Mimi codec reconstructs PCM audio from the 8 parallel audio token streams at the specified frame boundaries
Frequently Asked Questions
What is the audio frame duration in PersonaPlex?
Each audio frame represents 80 milliseconds of audio. This duration derives from the FRAME_RATE of 12.5 Hz defined in moshi/models/loaders.py#L40, calculated as 1/12.5 = 0.08 seconds. With a sample rate of 24 kHz, each frame contains exactly 1920 samples.
How many audio tokens does PersonaPlex generate per frame?
PersonaPlex generates 8 audio tokens per frame, one for each parallel codebook stream (dep_q=8). These tokens are produced alongside a single text token, resulting in a total of 9 tokens per generation step. The audio tokens are fed to the Mimi codec for reconstruction into PCM waveforms.
What are silence tokens used for in PersonaPlex?
Silence tokens are pre-defined token IDs ([948, 243, 1178, 546, 1736, 1030, 1978, 2008]) that encode a silent audio chunk when decoded by the Mimi model. Defined in moshi/models/lm.py#L56, these tokens allow the model to explicitly represent pauses or non-speech intervals in the audio stream without predicting raw audio samples.
Where is the zero-token defined and what does it do?
The zero-token is defined in moshi/models/lm.py via the zero_token_id() method (lines 379-380). This token reserves ID 0 to represent padding or "no sampling" positions in the attention matrix. It masks out unused positions during training loss computation (lines 549-551) and prevents the model from attending to or predicting content at placeholder locations in the sequence.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →