# PersonaPlex Voice Embedding File Formats: Complete Guide to Audio and Checkpoint Support

> Explore PersonaPlex voice embedding file formats. Learn about raw audio (WAV, FLAC, MP3) and PyTorch checkpoint (.pt) file support for seamless voice integration. Get the complete guide.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: api-reference
- Published: 2026-04-07

---

**PersonaPlex supports two distinct voice embedding file formats: raw audio files (WAV, FLAC, MP3) that are encoded on-the-fly via audio processing pipelines, and PyTorch checkpoint files (.pt) containing pre-computed embeddings with streaming cache states.**

NVIDIA's PersonaPlex voice generation system accepts both raw audio inputs and serialized embedding checkpoints for voice conditioning. Understanding these supported file formats is essential for optimizing inference latency and voice quality in production deployments, whether you need real-time voice cloning or efficient reuse of processed speaker characteristics.

## Raw Audio vs. Pre-Computed Embeddings: The Two Supported Formats

PersonaPlex handles voice conditioning through two distinct pathways defined in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) and [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py).

### Raw Voice Prompts (Audio Files)

Any audio format compatible with `sphn.read`—including **.wav**, **.flac**, and **.mp3**—serves as valid input for on-the-fly voice encoding. When you provide a raw audio path, the `LMGen.load_voice_prompt()` method processes the file through `load_audio()`, which internally invokes `sphn.read` to handle resampling and normalization before encoding.

In [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) (lines 60-66), the audio loading pipeline reads and prepares these files for the language model's voice conditioning mechanism.

### Pre-Computed Embedding Checkpoints (.pt Files)

For instant voice cloning without runtime encoding overhead, PersonaPlex accepts **PyTorch checkpoint files** with the `.pt` extension. These checkpoints contain serialized tensors stored under the keys `"embeddings"` and `"cache"`, representing previously processed voice data and associated streaming states.

The `LMGen.load_voice_prompt_embeddings()` method (defined in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py), lines 77-84) restores these tensors using `torch.load()`, moving them directly to the target device.

## Automatic File Format Detection

The PersonaPlex server ([`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py), lines 164-168) automatically routes requests based on file extension:

```python
if voice_prompt_path.endswith('.pt'):
    self.lm_gen.load_voice_prompt_embeddings(voice_prompt_path)
else:
    self.lm_gen.load_voice_prompt(voice_prompt_path)

```

This logic ensures **.pt** files trigger the embedding loader, while any other extension routes through the audio pipeline.

## Saving Voice Embeddings for Reuse

To generate reusable `.pt` checkpoints from raw audio, enable the `--save-voice-prompt-embeddings` flag. The saving mechanism (located in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py), lines 55-61) serializes both the stacked embeddings and streaming cache:

```python
torch.save(
    {
        "embeddings": torch.stack(saved_embeddings, dim=0).detach().cpu(),
        "cache": self._streaming_state.cache,
    },
    splitext(self.voice_prompt)[0] + ".pt",
)

```

## Practical Implementation Examples

### Loading Raw Audio Voice Prompts

Import `LMGen` and process a WAV file through the audio pipeline:

```python
from moshi.models.lm import LMGen

lm = LMGen(...)
lm.load_voice_prompt("assets/voice_prompts/speaker_voice.wav")

```

The audio undergoes resampling and normalization via `sphn.read` before voice encoding occurs.

### Loading Pre-Computed Embedding Checkpoints

For production deployments requiring minimal latency, load cached embeddings directly:

```python
from moshi.models.lm import LMGen

lm = LMGen(...)
lm.load_voice_prompt_embeddings("assets/voice_prompts/speaker_voice.pt")

```

This bypasses audio preprocessing and immediately applies the stored voice characteristics.

### Command-Line Interface Usage

Specify either format using the `--voice-prompt` argument in the offline inference script:

**Pre-computed embeddings:**

```bash
python -m moshi.offline \
    --voice-prompt "speaker.pt" \
    --input-wav "input.wav" \
    --output-wav "output.wav"

```

**Raw audio processing:**

```bash
python -m moshi.offline \
    --voice-prompt "speaker.wav" \
    --input-wav "input.wav" \
    --output-wav "output.wav"

```

## Summary

- **Raw audio files** (WAV, FLAC, MP3) supported via `sphn.read` are processed on-the-fly through `LMGen.load_voice_prompt()` in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py)
- **PyTorch checkpoint files** (`.pt`) containing `"embeddings"` and `"cache"` tensors load instantly via `LMGen.load_voice_prompt_embeddings()`
- The server automatically detects format by checking for the `.pt` extension in [`moshi/moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/server.py)
- Embedding checkpoints are generated using `torch.save()` with the `--save-voice-prompt-embeddings` flag for reuse across sessions

## Frequently Asked Questions

### Can I use MP3 files directly with PersonaPlex without converting to WAV?

Yes. PersonaPlex accepts any audio format that the `sphn.read` library supports, which includes MP3, FLAC, and WAV. The `LMGen.load_voice_prompt()` method handles format detection, resampling, and normalization automatically when loading raw audio inputs.

### What is the performance difference between raw audio and .pt checkpoint files?

Pre-computed `.pt` checkpoints eliminate runtime audio encoding latency because the embeddings are loaded directly via `torch.load()`. Raw audio requires processing through the full encode pipeline including resampling in `load_audio()`, making `.pt` files significantly faster for repeated inference with the same voice.

### How do I create a reusable .pt voice embedding checkpoint?

Enable the `--save-voice-prompt-embeddings` flag when running inference with a raw audio file. The system saves a checkpoint containing the processed embeddings and streaming cache to a `.pt` file adjacent to your input audio, which can be loaded later via `load_voice_prompt_embeddings()`.

### What data is stored inside a PersonaPlex .pt checkpoint file?

Each checkpoint contains two PyTorch tensors: `"embeddings"` (a stacked tensor of voice representations) and `"cache"` (the associated streaming state). These are serialized using `torch.save()` in [`moshi/moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/moshi/models/lm.py) and restored exactly to the model's device when loaded.