Voice Prompt Conditioning in NVIDIA PersonaPlex: Implementation and Usage Guide
NVIDIA PersonaPlex conditions its generative language-audio model on a speaker's voice prompt through a three-stage pipeline that loads audio into the LMGen class, encodes frames via the Mimi codec into KV cache representations, and pre-injects this conditioning before processing user audio.
NVIDIA PersonaPlex is an open-source voice interaction system that uses voice prompt conditioning to achieve zero-shot speaker consistency. The repository implements this capability through the LMGen class in moshi/moshi/models/lm.py, which orchestrates the loading, encoding, and injection of voice prompts. This article examines the technical implementation across the codebase to demonstrate how PersonaPlex prepares and utilizes voice prompts for low-latency, speaker-specific generation.
The Three-Stage Voice Prompt Conditioning Pipeline
PersonaPlex performs voice prompt conditioning through three tightly coupled stages that execute before any user interaction begins.
Stage 1: Voice-Prompt Loading and Normalization
The pipeline begins when LMGen.load_voice_prompt() reads a raw WAV file or pre-computed embedding file. Located in moshi/moshi/models/lm.py, this method normalizes the audio to -24 LUFS and stores it in the LMGen.voice_prompt_audio attribute. For pre-computed embeddings, LMGen.load_voice_prompt_embeddings() loads .pt files directly into LMGen.voice_prompt_embeddings.
Stage 2: Embedding and KV Cache Preparation
If processing raw audio, the system encodes the prompt frame-by-frame through the Mimi codec. The core logic resides in LMGen._step_voice_prompt_core() and LMGen._step_voice_prompt_frame(). During encoding, each frame's token sequence is fed to the language model with a zero text token—a special "system" token that ensures only the audio stream influences the internal KV cache. This creates a cached representation of the speaker's voice characteristics without text interference.
Stage 3: Prompt Injection into the Generation Loop
The prepared voice-prompt KV cache is pre-loaded into the model before the first user audio frame arrives. The LMGen.step_system_prompts() routine orchestrates this injection, running the voice prompt phase followed by a silence filler, an optional text prompt, and a final silence period. An asynchronous variant, step_system_prompts_async(), handles this non-blocking in server contexts. This guarantees the model's hidden state reflects the target speaker when streaming begins.
Implementation Across Source Files
The voice prompt conditioning logic spans three critical files in the NVIDIA/personaplex repository.
LMGen Class in lm.py
The moshi/moshi/models/lm.py file contains the core LMGen class that manages voice prompt state. Key methods include:
load_voice_prompt(): Handles RAW WAV ingestion and -24 LUFS normalization_step_voice_prompt_core(): Manages frame-wise encoding through the Mimi codec alongside_step_voice_prompt_frame()step_system_prompts(): Coordinates the full conditioning sequence including voice, silence, and text prompts
Server-Side Handling in server.py
For real-time applications, moshi/moshi/server.py implements the conditioning flow within the handle_chat() method. The server resolves the voice_prompt parameter from the HTTP request, calls self.lm_gen.load_voice_prompt() with the resolved path, and invokes await self.lm_gen.step_system_prompts_async() to condition the model before establishing the WebSocket audio stream.
Offline Inference in offline.py
Batch processing scenarios use moshi/moshi/offline.py, where the run_inference() function mirrors the server logic. It checks for .pt extensions to determine whether to call load_voice_prompt_embeddings() or load_voice_prompt(), then executes lm_gen.step_system_prompts() to inject prompts before processing user audio files.
Practical Code Examples
HTTP API Usage
Trigger voice prompt conditioning via the server's WebSocket endpoint by specifying the voice file in the query parameters:
curl "http://localhost:8080/chat?text_prompt=You+are+friendly&voice_prompt=alice.wav" \
-H "Connection: Upgrade" \
-H "Upgrade: websocket" \
--output -
The server resolves alice.wav within the directory configured by --voice-prompt-dir and executes the conditioning pipeline before streaming begins.
Python Offline Inference
Process audio files with custom voice prompts using the offline inference API:
from moshi.moshi.offline import run_inference
run_inference(
input_wav="user_query.wav",
output_wav="agent_reply.wav",
output_text="agent_reply.txt",
text_prompt="You are an enthusiastic tutor.",
voice_prompt_path="voices/teacher.wav",
hf_repo="NVIDIA/PersonaPlex",
device="cuda",
save_voice_prompt_embeddings=False,
)
This function automatically detects .pt embeddings versus raw audio and calls the appropriate loading method before generation.
Direct LMGen Integration
For custom implementations, instantiate LMGen and manually control the conditioning workflow:
from moshi.models import loaders, LMGen
# Load models and initialize generator
lm = loaders.get_moshi_lm("moshi.pt", device="cuda")
mimi = loaders.get_mimi("mimi.pt", device="cuda")
lm_gen = LMGen(
lm,
audio_silence_frame_cnt=24,
sample_rate=mimi.sample_rate,
device="cuda",
frame_rate=mimi.frame_rate
)
# Load and condition on voice prompt
lm_gen.load_voice_prompt("voices/jane.wav")
lm_gen.text_prompt_tokens = tokenizer.encode(
"<system> You are a helpful assistant. <system>"
)
lm_gen.step_system_prompts(mimi) # Injects voice + text prompts
This approach provides granular control over the 24-frame silence buffer and conditioning sequence.
Summary
- Voice prompt conditioning in PersonaPlex uses a three-stage pipeline: loading/normalization, Mimi codec encoding, and KV cache injection.
- The
LMGenclass inmoshi/moshi/models/lm.pymanages all conditioning logic through methods likeload_voice_prompt()andstep_system_prompts(). - Audio prompts are normalized to -24 LUFS and processed with zero text tokens to isolate speaker characteristics in the KV cache.
- Both server (
server.py) and offline (offline.py) implementations support raw WAV files or pre-computed.ptembeddings. - The conditioning occurs entirely before user audio processing, ensuring zero additional latency during generation.
Frequently Asked Questions
What audio format does PersonaPlex require for voice prompts?
PersonaPlex accepts standard WAV files for raw audio input. The LMGen.load_voice_prompt() method normalizes these to -24 LUFS before processing. Alternatively, you can provide pre-computed embeddings as .pt PyTorch files via load_voice_prompt_embeddings(), which bypasses the encoding step entirely.
How does voice prompt conditioning affect generation latency?
The conditioning pipeline executes entirely before the first user audio frame is processed. Since step_system_prompts() and step_system_prompts_async() complete before generation begins, voice prompt conditioning adds zero latency to the actual streaming inference, though it requires upfront computation time proportional to the prompt length.
Can I reuse pre-computed voice prompt embeddings?
Yes. PersonaPlex supports saving and loading pre-computed embeddings through LMGen.load_voice_prompt_embeddings(). When provided with a .pt file path in offline.py or via the API, the system skips the Mimi codec encoding and loads the KV cache directly, significantly reducing startup time for repeated uses of the same speaker voice.
What is the role of the Mimi codec in voice prompt conditioning?
The Mimi codec compresses the raw voice prompt audio into discrete tokens frame-by-frame. During conditioning, LMGen processes these tokens through the language model with zero text tokens to populate the KV cache. This encoded representation captures the speaker's acoustic characteristics without requiring model fine-tuning, enabling zero-shot voice cloning capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →