# How SplitResidualVectorQuantizer Encodes Audio to Tokens in NVIDIA PersonaPlex

> Learn how SplitResidualVectorQuantizer encodes audio to tokens by dividing quantization between semantic and acoustic RVQs. Discover how this process prepares audio data for language models in NVIDIA PersonaPlex.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: internals
- Published: 2026-04-07

---

**The SplitResidualVectorQuantizer converts continuous audio embeddings into discrete token IDs by distributing the quantization workload across separate semantic and acoustic Residual Vector Quantizers, concatenating their outputs into a tensor of shape `[B, K, T]` suitable for language model consumption.**

The SplitResidualVectorQuantizer (SRVQ) serves as the core tokenization engine in NVIDIA's PersonaPlex (Moshi) architecture, bridging the gap between raw audio waveforms and discrete token sequences. Understanding how this component encodes audio to tokens is essential for working with the compression model and downstream speech-language models. This article examines the actual implementation in the NVIDIA/personaplex repository, tracing the exact code path from waveform input to discrete codebook indices.

## Audio Encoding Pipeline Overview

The encoding process traverses three distinct stages before producing discrete tokens. First, the **CompressionModel** encoder processes raw audio waveforms using Conv1D and transformer layers to produce continuous latent embeddings. Second, these embeddings pass through the SplitResidualVectorQuantizer's selection logic to determine active codebooks. Finally, the underlying Residual Vector Quantizers convert the continuous representations into discrete indices.

In [`moshi/moshi/models/compression.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/compression.py), the encoder produces a tensor `emb` with shape `[B, D, T']`, where `B` is batch size, `D` is embedding dimension, and `T'` represents the frame-rate-adjusted temporal length. The `_to_framerate` method resamples these embeddings to match the model's target frame rate before quantization.

## The SplitResidualVectorQuantizer.encode() Method

The primary entry point for audio-to-token conversion resides in [`moshi/moshi/quantization/vq.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/quantization/vq.py). The `encode` method implements a split design that separates high-level semantic information from low-level acoustic details across distinct codebook groups.

### Semantic Quantization with rvq_first

The encoding process begins with the semantic RVQ (`self.rvq_first`), which processes the input using `n_q_semantic` codebooks. This first quantizer captures high-level representations of the audio content.

```python

# moshi/moshi/quantization/vq.py – SplitResidualVectorQuantizer.encode

codes = self.rvq_first.encode(x)  # semantic RVQ

```

The semantic quantizer applies its own input projection layer before invoking the core vector quantization logic, allowing it to operate in a distinct feature space optimized for semantic content.

### Acoustic Quantization and Concatenation

When the total number of codebooks (`n_q`) exceeds the semantic allocation (`n_q_semantic`), the method activates the acoustic RVQ (`self.rvq_rest`) to encode residual information using the remaining codebooks. The `codebook_offset=1` parameter ensures these acoustic codebooks use a distinct index range, preventing collisions during concatenation.

```python
if self.n_q > self.n_q_semantic:  # need acoustic part?

    acoustic_codes = self.rvq_rest.encode(x)  # acoustic RVQ

    codes = torch.cat([codes, acoustic_codes], dim=1)

# → shape [B, K, T] where K = total active codebooks

```

This concatenation produces the final token tensor of shape `[B, K, T]`, where each temporal position `t` contains `K` discrete indices—one per active codebook.

## Underlying ResidualVectorQuantizer Implementation

Each RVQ instance encapsulates the actual quantization logic in its `encode` method. Before quantization, the input undergoes dimensionality adjustment through `input_proj`, followed by the core residual vector quantization process.

```python

# moshi/moshi/quantization/vq.py – ResidualVectorQuantizer.encode

x = self.input_proj(x)
codes = self.vq.encode(x, n_q=n_q)  # core VQ returns [K, B, T]

codes = codes.transpose(0, 1)       # → [B, K, T]

```

The core `ResidualVectorQuantization` class performs iterative residual quantization: it quantizes the input, computes the residual, and passes the residual to the next codebook. This process repeats for `n_q` iterations, encoding increasingly fine-grained details at each step.

## Complete Code Examples

The following examples demonstrate direct usage of the SplitResidualVectorQuantizer and the full compression pipeline.

**Direct SRVQ Usage:**

```python
import torch
from moshi.moshi.quantization.vq import SplitResidualVectorQuantizer

# Dummy audio batch: B=2, C=1 (mono), T=16000 (1 sec at 16 kHz)

audio = torch.randn(2, 1, 16000)

# Instantiate an SRVQ (8 total codebooks, 2 semantic)

srvq = SplitResidualVectorQuantizer(
    n_q=8,
    n_q_semantic=2,
    dimension=128,
    bins=1024,
    force_projection=True,
)

# Encode audio → token IDs

tokens = srvq.encode(audio)  # shape: [2, 8, 16000]

print(tokens.shape)         # torch.Size([2, 8, 16000])

```

**Full Compression Pipeline:**

```python
import torch
from moshi.moshi.models.compression import CompressionModel

# Load a pretrained compression checkpoint

model = CompressionModel.from_pretrained("moshi_base")
model.eval()

# Input waveform: 2 seconds at 16kHz

waveform = torch.randn(1, 1, 32000)

# Encode to discrete tokens

tokens = model.encode(waveform)  # uses SplitResidualVectorQuantizer internally

print(tokens.shape)             # e.g. torch.Size([1, 8, 32000])

```

## Summary

- **SplitResidualVectorQuantizer** in [`moshi/moshi/quantization/vq.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/quantization/vq.py) serves as the primary audio-to-token encoder in NVIDIA PersonaPlex.
- The implementation splits quantization across **semantic** (`rvq_first`) and **acoustic** (`rvq_rest`) Residual Vector Quantizers.
- The `encode` method concatenates semantic and acoustic codes along the codebook dimension, producing output of shape `[B, K, T]`.
- Each underlying **ResidualVectorQuantizer** applies input projection before calling the core `ResidualVectorQuantization` engine.
- The resulting tokens are discrete indices suitable for direct consumption by transformer-based language models.

## Frequently Asked Questions

### What is the difference between semantic and acoustic codebooks in SRVQ?

The semantic codebooks (`n_q_semantic`) capture high-level audio content and prosody, while the acoustic codebooks handle fine-grained waveform reconstruction details. According to the NVIDIA/personaplex source code, the semantic RVQ (`rvq_first`) operates independently with its own projection layer, allowing it to optimize for abstract representations, whereas the acoustic RVQ (`rvq_rest`) processes the same input with `codebook_offset=1` to encode residual information using distinct codebook indices.

### How does the shape of the output tensor change during encoding?

The core vector quantization returns indices of shape `[K, B, T]`, which the `ResidualVectorQuantizer.encode` method transposes to `[B, K, T]` for batch-first convention. When SplitResidualVectorQuantizer concatenates semantic and acoustic codes, it combines them along dimension 1 (the codebook dimension), maintaining the `[B, K, T]` structure where `K` equals the total active codebooks (`n_q`).

### Can I use SplitResidualVectorQuantizer independently of the CompressionModel?

Yes, the SRVQ class can be instantiated and used directly on continuous embeddings, as demonstrated in the direct usage example. However, for raw audio waveforms, you must first extract latent embeddings using the encoder stack from `CompressionModel`, typically accessed through `model.encode()` which handles the waveform preprocessing, encoder forward pass, and framerate adjustment before calling the quantizer.

### What determines the number of active codebooks during encoding?

The `n_q` parameter controls how many codebooks are active during the forward pass. If `n_q` is less than or equal to `n_q_semantic`, only the semantic RVQ executes. When `n_q` exceeds `n_q_semantic`, the acoustic RVQ activates for the remaining codebook slots. This allows dynamic adjustment of the bit rate and reconstruction quality by varying the number of residual quantization steps.