# Voice Prompt Conditioning in NVIDIA PersonaPlex: Implementation and Usage Guide

> Learn how to implement voice prompt conditioning in NVIDIA PersonaPlex. This guide details the three-stage pipeline for loading, encoding, and pre-injecting speaker audio into the generative model.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: how-to-guide
- Published: 2026-04-07

---

**NVIDIA PersonaPlex conditions its generative language-audio model on a speaker's voice prompt through a three-stage pipeline that loads audio into the `LMGen` class, encodes frames via the Mimi codec into KV cache representations, and pre-injects this conditioning before processing user audio.**

NVIDIA PersonaPlex is an open-source voice interaction system that uses voice prompt conditioning to achieve zero-shot speaker consistency. The repository implements this capability through the `LMGen` class in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py), which orchestrates the loading, encoding, and injection of voice prompts. This article examines the technical implementation across the codebase to demonstrate how PersonaPlex prepares and utilizes voice prompts for low-latency, speaker-specific generation.

## The Three-Stage Voice Prompt Conditioning Pipeline

PersonaPlex performs voice prompt conditioning through three tightly coupled stages that execute before any user interaction begins.

### Stage 1: Voice-Prompt Loading and Normalization

The pipeline begins when `LMGen.load_voice_prompt()` reads a raw WAV file or pre-computed embedding file. Located in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py), this method normalizes the audio to **-24 LUFS** and stores it in the `LMGen.voice_prompt_audio` attribute. For pre-computed embeddings, `LMGen.load_voice_prompt_embeddings()` loads `.pt` files directly into `LMGen.voice_prompt_embeddings`.

### Stage 2: Embedding and KV Cache Preparation

If processing raw audio, the system encodes the prompt frame-by-frame through the **Mimi** codec. The core logic resides in `LMGen._step_voice_prompt_core()` and `LMGen._step_voice_prompt_frame()`. During encoding, each frame's token sequence is fed to the language model with a **zero text token**—a special "system" token that ensures only the audio stream influences the internal KV cache. This creates a cached representation of the speaker's voice characteristics without text interference.

### Stage 3: Prompt Injection into the Generation Loop

The prepared voice-prompt KV cache is pre-loaded into the model before the first user audio frame arrives. The `LMGen.step_system_prompts()` routine orchestrates this injection, running the voice prompt phase followed by a silence filler, an optional text prompt, and a final silence period. An asynchronous variant, `step_system_prompts_async()`, handles this non-blocking in server contexts. This guarantees the model's hidden state reflects the target speaker when streaming begins.

## Implementation Across Source Files

The voice prompt conditioning logic spans three critical files in the NVIDIA/personaplex repository.

### LMGen Class in lm.py

The [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) file contains the core `LMGen` class that manages voice prompt state. Key methods include:

- **`load_voice_prompt()`**: Handles RAW WAV ingestion and -24 LUFS normalization
- **`_step_voice_prompt_core()`**: Manages frame-wise encoding through the Mimi codec alongside `_step_voice_prompt_frame()`
- **`step_system_prompts()`**: Coordinates the full conditioning sequence including voice, silence, and text prompts

### Server-Side Handling in server.py

For real-time applications, [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py) implements the conditioning flow within the `handle_chat()` method. The server resolves the `voice_prompt` parameter from the HTTP request, calls `self.lm_gen.load_voice_prompt()` with the resolved path, and invokes `await self.lm_gen.step_system_prompts_async()` to condition the model before establishing the WebSocket audio stream.

### Offline Inference in offline.py

Batch processing scenarios use [`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py), where the `run_inference()` function mirrors the server logic. It checks for `.pt` extensions to determine whether to call `load_voice_prompt_embeddings()` or `load_voice_prompt()`, then executes `lm_gen.step_system_prompts()` to inject prompts before processing user audio files.

## Practical Code Examples

### HTTP API Usage

Trigger voice prompt conditioning via the server's WebSocket endpoint by specifying the voice file in the query parameters:

```bash
curl "http://localhost:8080/chat?text_prompt=You+are+friendly&voice_prompt=alice.wav" \
     -H "Connection: Upgrade" \
     -H "Upgrade: websocket" \
     --output -

```

The server resolves `alice.wav` within the directory configured by `--voice-prompt-dir` and executes the conditioning pipeline before streaming begins.

### Python Offline Inference

Process audio files with custom voice prompts using the offline inference API:

```python
from moshi.moshi.offline import run_inference

run_inference(
    input_wav="user_query.wav",
    output_wav="agent_reply.wav",
    output_text="agent_reply.txt",
    text_prompt="You are an enthusiastic tutor.",
    voice_prompt_path="voices/teacher.wav",
    hf_repo="NVIDIA/PersonaPlex",
    device="cuda",
    save_voice_prompt_embeddings=False,
)

```

This function automatically detects `.pt` embeddings versus raw audio and calls the appropriate loading method before generation.

### Direct LMGen Integration

For custom implementations, instantiate `LMGen` and manually control the conditioning workflow:

```python
from moshi.models import loaders, LMGen

# Load models and initialize generator

lm = loaders.get_moshi_lm("moshi.pt", device="cuda")
mimi = loaders.get_mimi("mimi.pt", device="cuda")
lm_gen = LMGen(
    lm, 
    audio_silence_frame_cnt=24,
    sample_rate=mimi.sample_rate,
    device="cuda",
    frame_rate=mimi.frame_rate
)

# Load and condition on voice prompt

lm_gen.load_voice_prompt("voices/jane.wav")
lm_gen.text_prompt_tokens = tokenizer.encode(
    "<system> You are a helpful assistant. <system>"
)
lm_gen.step_system_prompts(mimi)  # Injects voice + text prompts

```

This approach provides granular control over the **24-frame silence buffer** and conditioning sequence.

## Summary

- **Voice prompt conditioning** in PersonaPlex uses a three-stage pipeline: loading/normalization, Mimi codec encoding, and KV cache injection.
- The `LMGen` class in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) manages all conditioning logic through methods like `load_voice_prompt()` and `step_system_prompts()`.
- Audio prompts are normalized to **-24 LUFS** and processed with **zero text tokens** to isolate speaker characteristics in the KV cache.
- Both server ([`server.py`](https://github.com/NVIDIA/personaplex/blob/main/server.py)) and offline ([`offline.py`](https://github.com/NVIDIA/personaplex/blob/main/offline.py)) implementations support raw WAV files or pre-computed `.pt` embeddings.
- The conditioning occurs entirely before user audio processing, ensuring zero additional latency during generation.

## Frequently Asked Questions

### What audio format does PersonaPlex require for voice prompts?

PersonaPlex accepts standard WAV files for raw audio input. The `LMGen.load_voice_prompt()` method normalizes these to -24 LUFS before processing. Alternatively, you can provide pre-computed embeddings as `.pt` PyTorch files via `load_voice_prompt_embeddings()`, which bypasses the encoding step entirely.

### How does voice prompt conditioning affect generation latency?

The conditioning pipeline executes entirely before the first user audio frame is processed. Since `step_system_prompts()` and `step_system_prompts_async()` complete before generation begins, voice prompt conditioning adds **zero latency** to the actual streaming inference, though it requires upfront computation time proportional to the prompt length.

### Can I reuse pre-computed voice prompt embeddings?

Yes. PersonaPlex supports saving and loading pre-computed embeddings through `LMGen.load_voice_prompt_embeddings()`. When provided with a `.pt` file path in [`offline.py`](https://github.com/NVIDIA/personaplex/blob/main/offline.py) or via the API, the system skips the Mimi codec encoding and loads the KV cache directly, significantly reducing startup time for repeated uses of the same speaker voice.

### What is the role of the Mimi codec in voice prompt conditioning?

The **Mimi** codec compresses the raw voice prompt audio into discrete tokens frame-by-frame. During conditioning, `LMGen` processes these tokens through the language model with zero text tokens to populate the KV cache. This encoded representation captures the speaker's acoustic characteristics without requiring model fine-tuning, enabling zero-shot voice cloning capabilities.