# How to Extract Voice Embeddings from WAV Files in PersonaPlex

> Extract voice embeddings from WAV files in PersonaPlex by enabling save_voice_prompt_embeddings True. Cache embeddings as PyTorch tensors for faster inference.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: how-to-guide
- Published: 2026-04-07

---

**Enable `save_voice_prompt_embeddings=True` in the `LMGen` class to capture per-frame voice embeddings from WAV files, then cache them as PyTorch tensors in a `.pt` file for efficient reuse across inference sessions.**

NVIDIA's PersonaPlex repository implements the **Moshi** voice interaction model, which processes spoken prompts through a Mimi audio encoder before projecting them through language model embedding tables. Extracting voice embeddings from WAV files allows you to serialize the acoustic representation of a speaker—avoiding redundant computation when reproducing the same voice persona across multiple generations.

## Understanding the Voice Embedding Architecture

The extraction workflow in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) follows a deterministic three-stage pipeline.

### Stage 1: Audio Loading and Normalization

The `LMGen.load_voice_prompt()` method handles ingestion of raw WAV data. It invokes `load_audio` to normalize the input to **-24 LUFS** and resample to mono-channel, storing the result in `self.voice_prompt_audio` ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L60-L74)).

### Stage 2: Frame-wise Encoding

The `_encode_voice_prompt_frames` generator iterates over the normalized audio using `_iterate_audio`, applying `encode_from_sphn` to produce **Moshi token tensors** for each temporal slice. This yields a sequence of encoded frames ready for language model consumption ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L99-L108)).

### Stage 3: Embedding Capture and Persistence

When `save_voice_prompt_embeddings=True` is passed to the `LMGen` constructor, `_step_voice_prompt_core` accumulates hidden states during the forward pass. The method returns `(output, embeddings)` tuples, which are aggregated in `saved_embeddings` and ultimately serialized via `torch.save` ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L110-L124)).

## Command-Line Extraction

For batch processing or offline workflows, use the [`offline.py`](https://github.com/NVIDIA/personaplex/blob/main/offline.py) script bundled in the repository. The `--save-voice-prompt-embeddings` flag triggers the embedding capture logic during voice prompt processing.

```bash
python -m moshi.moshi.offline \
    --input-wav assets/test/input_service.wav \
    --output-wav out.wav \
    --output-text out.txt \
    --text-prompt "" \
    --voice-prompt my_prompt.wav \
    --save-voice-prompt-embeddings

```

The script logic in [`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py) automatically detects file extensions to determine whether to load a raw WAV or a pre-computed checkpoint. When processing raw audio, it instantiates `LMGen` with `save_voice_prompt_embeddings=True`, executes the frame-wise encoding loop, and writes a `my_prompt.pt` file containing both the `embeddings` tensor and the model cache state ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py#L34-L41)).

## Python API for Custom Workflows

For programmatic control, instantiate `LMGen` directly with the persistence flag enabled, then drive the voice prompt encoder manually.

```python
import torch
from moshi.moshi.models.lm import LMGen
from moshi.moshi import loaders

# Initialize model components

device = "cuda"
mimi = loaders.get_mimi(loaders.MIMI_NAME, device=device)
lm = loaders.get_moshi_lm(loaders.MOSHI_NAME, device=device)

# Configure generator to save embeddings

gen = LMGen(
    lm,
    audio_silence_frame_cnt=int(0.5 * mimi.frame_rate),
    sample_rate=mimi.sample_rate,
    device=device,
    frame_rate=mimi.frame_rate,
    save_voice_prompt_embeddings=True,  # Enable extraction

)

# Load and encode voice prompt

gen.load_voice_prompt("speaker_voice.wav")

# Drive the encoding loop to completion

for _ in gen._step_voice_prompt_core(mimi):
    pass

# Access embeddings: shape [num_frames, embedding_dim]

embeddings = gen.voice_prompt_embeddings

# Persist for reuse

torch.save(
    {"embeddings": embeddings.cpu(), "cache": gen._streaming_state.cache.cpu()},
    "speaker_cache.pt"
)

```

The `load_voice_prompt` method normalizes the input according to the Mimi codec requirements, while `_step_voice_prompt_core` handles the actual inference steps that populate `voice_prompt_embeddings` when the constructor flag is active ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L125-L160)).

## Reusing Cached Voice Embeddings

Once extracted, voice embeddings eliminate redundant computation. Load a `.pt` checkpoint directly via `load_voice_prompt_embeddings` rather than reprocessing the raw WAV.

```python

# Bypass encoding entirely

gen.load_voice_prompt_embeddings("speaker_cache.pt")

# Embeddings now available in gen.voice_prompt_embeddings

```

This method restores both the embedding tensor and the transformer cache state required for consistent voice cloning, as implemented in [`lm.py`](https://github.com/NVIDIA/personaplex/blob/main/lm.py) lines 77-84 ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L77-L84)).

## Summary

- **Enable persistence**: Set `save_voice_prompt_embeddings=True` in `LMGen` to activate the embedding capture logic defined in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py).
- **Process stages**: The pipeline moves through `load_voice_prompt` (normalization), `_encode_voice_prompt_frames` (tokenization), and `_step_voice_prompt_core` (embedding extraction).
- **Cache format**: Embeddings are stored as `[num_frames, embedding_dim]` tensors alongside model `cache` states in `.pt` checkpoints compatible with `torch.load`.
- **Reuse strategy**: Check for existing `.pt` files before processing raw WAVs to minimize latency; use `load_voice_prompt_embeddings` to restore saved states instantly.

## Frequently Asked Questions

### What is the tensor shape of extracted voice embeddings?

The `voice_prompt_embeddings` attribute contains a PyTorch tensor with shape `[num_frames, embedding_dim]`, where `num_frames` corresponds to the temporal duration of the audio after Mimi encoding, and `embedding_dim` matches the hidden dimension of the Moshi language model.

### How does PersonaPlex normalize incoming WAV files?

The `load_voice_prompt` method in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) applies ITU-R BS.1770-4 loudness normalization to **-24 LUFS** and converts multi-channel audio to mono, ensuring consistent scaling across varying input sources ([source](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py#L60-L74)).

### Can I extract embeddings without running full audio generation?

Yes. The `_step_voice_prompt_core` method can be driven to completion independently of the main audio generation loop. Simply instantiate `LMGen`, load your voice prompt, and iterate through the core stepping loop—the embeddings populate in `voice_prompt_embeddings` after processing the final frame without requiring token generation.

### Where does the embedding extraction logic reside in the source tree?

Core functionality lives in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) within the `LMGen` class, specifically lines 60-160 covering `load_voice_prompt`, `_encode_voice_prompt_frames`, and `_step_voice_prompt_core`. The CLI wrapper implementing `--save-voice-prompt-embeddings` resides in [`moshi/moshi/offline.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/offline.py).