How to Extract Voice Embeddings from WAV Files in PersonaPlex

Enable save_voice_prompt_embeddings=True in the LMGen class to capture per-frame voice embeddings from WAV files, then cache them as PyTorch tensors in a .pt file for efficient reuse across inference sessions.

NVIDIA's PersonaPlex repository implements the Moshi voice interaction model, which processes spoken prompts through a Mimi audio encoder before projecting them through language model embedding tables. Extracting voice embeddings from WAV files allows you to serialize the acoustic representation of a speaker—avoiding redundant computation when reproducing the same voice persona across multiple generations.

Understanding the Voice Embedding Architecture

The extraction workflow in moshi/moshi/models/lm.py follows a deterministic three-stage pipeline.

Stage 1: Audio Loading and Normalization

The LMGen.load_voice_prompt() method handles ingestion of raw WAV data. It invokes load_audio to normalize the input to -24 LUFS and resample to mono-channel, storing the result in self.voice_prompt_audio (source).

Stage 2: Frame-wise Encoding

The _encode_voice_prompt_frames generator iterates over the normalized audio using _iterate_audio, applying encode_from_sphn to produce Moshi token tensors for each temporal slice. This yields a sequence of encoded frames ready for language model consumption (source).

Stage 3: Embedding Capture and Persistence

When save_voice_prompt_embeddings=True is passed to the LMGen constructor, _step_voice_prompt_core accumulates hidden states during the forward pass. The method returns (output, embeddings) tuples, which are aggregated in saved_embeddings and ultimately serialized via torch.save (source).

Command-Line Extraction

For batch processing or offline workflows, use the offline.py script bundled in the repository. The --save-voice-prompt-embeddings flag triggers the embedding capture logic during voice prompt processing.

python -m moshi.moshi.offline \
    --input-wav assets/test/input_service.wav \
    --output-wav out.wav \
    --output-text out.txt \
    --text-prompt "" \
    --voice-prompt my_prompt.wav \
    --save-voice-prompt-embeddings

The script logic in moshi/moshi/offline.py automatically detects file extensions to determine whether to load a raw WAV or a pre-computed checkpoint. When processing raw audio, it instantiates LMGen with save_voice_prompt_embeddings=True, executes the frame-wise encoding loop, and writes a my_prompt.pt file containing both the embeddings tensor and the model cache state (source).

Python API for Custom Workflows

For programmatic control, instantiate LMGen directly with the persistence flag enabled, then drive the voice prompt encoder manually.

import torch
from moshi.moshi.models.lm import LMGen
from moshi.moshi import loaders

# Initialize model components

device = "cuda"
mimi = loaders.get_mimi(loaders.MIMI_NAME, device=device)
lm = loaders.get_moshi_lm(loaders.MOSHI_NAME, device=device)

# Configure generator to save embeddings

gen = LMGen(
    lm,
    audio_silence_frame_cnt=int(0.5 * mimi.frame_rate),
    sample_rate=mimi.sample_rate,
    device=device,
    frame_rate=mimi.frame_rate,
    save_voice_prompt_embeddings=True,  # Enable extraction

)

# Load and encode voice prompt

gen.load_voice_prompt("speaker_voice.wav")

# Drive the encoding loop to completion

for _ in gen._step_voice_prompt_core(mimi):
    pass

# Access embeddings: shape [num_frames, embedding_dim]

embeddings = gen.voice_prompt_embeddings

# Persist for reuse

torch.save(
    {"embeddings": embeddings.cpu(), "cache": gen._streaming_state.cache.cpu()},
    "speaker_cache.pt"
)

The load_voice_prompt method normalizes the input according to the Mimi codec requirements, while _step_voice_prompt_core handles the actual inference steps that populate voice_prompt_embeddings when the constructor flag is active (source).

Reusing Cached Voice Embeddings

Once extracted, voice embeddings eliminate redundant computation. Load a .pt checkpoint directly via load_voice_prompt_embeddings rather than reprocessing the raw WAV.


# Bypass encoding entirely

gen.load_voice_prompt_embeddings("speaker_cache.pt")

# Embeddings now available in gen.voice_prompt_embeddings

This method restores both the embedding tensor and the transformer cache state required for consistent voice cloning, as implemented in lm.py lines 77-84 (source).

Summary

  • Enable persistence: Set save_voice_prompt_embeddings=True in LMGen to activate the embedding capture logic defined in moshi/moshi/models/lm.py.
  • Process stages: The pipeline moves through load_voice_prompt (normalization), _encode_voice_prompt_frames (tokenization), and _step_voice_prompt_core (embedding extraction).
  • Cache format: Embeddings are stored as [num_frames, embedding_dim] tensors alongside model cache states in .pt checkpoints compatible with torch.load.
  • Reuse strategy: Check for existing .pt files before processing raw WAVs to minimize latency; use load_voice_prompt_embeddings to restore saved states instantly.

Frequently Asked Questions

What is the tensor shape of extracted voice embeddings?

The voice_prompt_embeddings attribute contains a PyTorch tensor with shape [num_frames, embedding_dim], where num_frames corresponds to the temporal duration of the audio after Mimi encoding, and embedding_dim matches the hidden dimension of the Moshi language model.

How does PersonaPlex normalize incoming WAV files?

The load_voice_prompt method in moshi/moshi/models/lm.py applies ITU-R BS.1770-4 loudness normalization to -24 LUFS and converts multi-channel audio to mono, ensuring consistent scaling across varying input sources (source).

Can I extract embeddings without running full audio generation?

Yes. The _step_voice_prompt_core method can be driven to completion independently of the main audio generation loop. Simply instantiate LMGen, load your voice prompt, and iterate through the core stepping loop—the embeddings populate in voice_prompt_embeddings after processing the final frame without requiring token generation.

Where does the embedding extraction logic reside in the source tree?

Core functionality lives in moshi/moshi/models/lm.py within the LMGen class, specifically lines 60-160 covering load_voice_prompt, _encode_voice_prompt_frames, and _step_voice_prompt_core. The CLI wrapper implementing --save-voice-prompt-embeddings resides in moshi/moshi/offline.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →