How to Extract Voice Embeddings from WAV Files in PersonaPlex
Enable save_voice_prompt_embeddings=True in the LMGen class to capture per-frame voice embeddings from WAV files, then cache them as PyTorch tensors in a .pt file for efficient reuse across inference sessions.
NVIDIA's PersonaPlex repository implements the Moshi voice interaction model, which processes spoken prompts through a Mimi audio encoder before projecting them through language model embedding tables. Extracting voice embeddings from WAV files allows you to serialize the acoustic representation of a speaker—avoiding redundant computation when reproducing the same voice persona across multiple generations.
Understanding the Voice Embedding Architecture
The extraction workflow in moshi/moshi/models/lm.py follows a deterministic three-stage pipeline.
Stage 1: Audio Loading and Normalization
The LMGen.load_voice_prompt() method handles ingestion of raw WAV data. It invokes load_audio to normalize the input to -24 LUFS and resample to mono-channel, storing the result in self.voice_prompt_audio (source).
Stage 2: Frame-wise Encoding
The _encode_voice_prompt_frames generator iterates over the normalized audio using _iterate_audio, applying encode_from_sphn to produce Moshi token tensors for each temporal slice. This yields a sequence of encoded frames ready for language model consumption (source).
Stage 3: Embedding Capture and Persistence
When save_voice_prompt_embeddings=True is passed to the LMGen constructor, _step_voice_prompt_core accumulates hidden states during the forward pass. The method returns (output, embeddings) tuples, which are aggregated in saved_embeddings and ultimately serialized via torch.save (source).
Command-Line Extraction
For batch processing or offline workflows, use the offline.py script bundled in the repository. The --save-voice-prompt-embeddings flag triggers the embedding capture logic during voice prompt processing.
python -m moshi.moshi.offline \
--input-wav assets/test/input_service.wav \
--output-wav out.wav \
--output-text out.txt \
--text-prompt "" \
--voice-prompt my_prompt.wav \
--save-voice-prompt-embeddings
The script logic in moshi/moshi/offline.py automatically detects file extensions to determine whether to load a raw WAV or a pre-computed checkpoint. When processing raw audio, it instantiates LMGen with save_voice_prompt_embeddings=True, executes the frame-wise encoding loop, and writes a my_prompt.pt file containing both the embeddings tensor and the model cache state (source).
Python API for Custom Workflows
For programmatic control, instantiate LMGen directly with the persistence flag enabled, then drive the voice prompt encoder manually.
import torch
from moshi.moshi.models.lm import LMGen
from moshi.moshi import loaders
# Initialize model components
device = "cuda"
mimi = loaders.get_mimi(loaders.MIMI_NAME, device=device)
lm = loaders.get_moshi_lm(loaders.MOSHI_NAME, device=device)
# Configure generator to save embeddings
gen = LMGen(
lm,
audio_silence_frame_cnt=int(0.5 * mimi.frame_rate),
sample_rate=mimi.sample_rate,
device=device,
frame_rate=mimi.frame_rate,
save_voice_prompt_embeddings=True, # Enable extraction
)
# Load and encode voice prompt
gen.load_voice_prompt("speaker_voice.wav")
# Drive the encoding loop to completion
for _ in gen._step_voice_prompt_core(mimi):
pass
# Access embeddings: shape [num_frames, embedding_dim]
embeddings = gen.voice_prompt_embeddings
# Persist for reuse
torch.save(
{"embeddings": embeddings.cpu(), "cache": gen._streaming_state.cache.cpu()},
"speaker_cache.pt"
)
The load_voice_prompt method normalizes the input according to the Mimi codec requirements, while _step_voice_prompt_core handles the actual inference steps that populate voice_prompt_embeddings when the constructor flag is active (source).
Reusing Cached Voice Embeddings
Once extracted, voice embeddings eliminate redundant computation. Load a .pt checkpoint directly via load_voice_prompt_embeddings rather than reprocessing the raw WAV.
# Bypass encoding entirely
gen.load_voice_prompt_embeddings("speaker_cache.pt")
# Embeddings now available in gen.voice_prompt_embeddings
This method restores both the embedding tensor and the transformer cache state required for consistent voice cloning, as implemented in lm.py lines 77-84 (source).
Summary
- Enable persistence: Set
save_voice_prompt_embeddings=TrueinLMGento activate the embedding capture logic defined inmoshi/moshi/models/lm.py. - Process stages: The pipeline moves through
load_voice_prompt(normalization),_encode_voice_prompt_frames(tokenization), and_step_voice_prompt_core(embedding extraction). - Cache format: Embeddings are stored as
[num_frames, embedding_dim]tensors alongside modelcachestates in.ptcheckpoints compatible withtorch.load. - Reuse strategy: Check for existing
.ptfiles before processing raw WAVs to minimize latency; useload_voice_prompt_embeddingsto restore saved states instantly.
Frequently Asked Questions
What is the tensor shape of extracted voice embeddings?
The voice_prompt_embeddings attribute contains a PyTorch tensor with shape [num_frames, embedding_dim], where num_frames corresponds to the temporal duration of the audio after Mimi encoding, and embedding_dim matches the hidden dimension of the Moshi language model.
How does PersonaPlex normalize incoming WAV files?
The load_voice_prompt method in moshi/moshi/models/lm.py applies ITU-R BS.1770-4 loudness normalization to -24 LUFS and converts multi-channel audio to mono, ensuring consistent scaling across varying input sources (source).
Can I extract embeddings without running full audio generation?
Yes. The _step_voice_prompt_core method can be driven to completion independently of the main audio generation loop. Simply instantiate LMGen, load your voice prompt, and iterate through the core stepping loop—the embeddings populate in voice_prompt_embeddings after processing the final frame without requiring token generation.
Where does the embedding extraction logic reside in the source tree?
Core functionality lives in moshi/moshi/models/lm.py within the LMGen class, specifically lines 60-160 covering load_voice_prompt, _encode_voice_prompt_frames, and _step_voice_prompt_core. The CLI wrapper implementing --save-voice-prompt-embeddings resides in moshi/moshi/offline.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →