PersonaPlex Voice Embedding File Formats: Complete Guide to Audio and Checkpoint Support
PersonaPlex supports two distinct voice embedding file formats: raw audio files (WAV, FLAC, MP3) that are encoded on-the-fly via audio processing pipelines, and PyTorch checkpoint files (.pt) containing pre-computed embeddings with streaming cache states.
NVIDIA's PersonaPlex voice generation system accepts both raw audio inputs and serialized embedding checkpoints for voice conditioning. Understanding these supported file formats is essential for optimizing inference latency and voice quality in production deployments, whether you need real-time voice cloning or efficient reuse of processed speaker characteristics.
Raw Audio vs. Pre-Computed Embeddings: The Two Supported Formats
PersonaPlex handles voice conditioning through two distinct pathways defined in moshi/moshi/models/lm.py and moshi/moshi/server.py.
Raw Voice Prompts (Audio Files)
Any audio format compatible with sphn.read—including .wav, .flac, and .mp3—serves as valid input for on-the-fly voice encoding. When you provide a raw audio path, the LMGen.load_voice_prompt() method processes the file through load_audio(), which internally invokes sphn.read to handle resampling and normalization before encoding.
In moshi/moshi/models/lm.py (lines 60-66), the audio loading pipeline reads and prepares these files for the language model's voice conditioning mechanism.
Pre-Computed Embedding Checkpoints (.pt Files)
For instant voice cloning without runtime encoding overhead, PersonaPlex accepts PyTorch checkpoint files with the .pt extension. These checkpoints contain serialized tensors stored under the keys "embeddings" and "cache", representing previously processed voice data and associated streaming states.
The LMGen.load_voice_prompt_embeddings() method (defined in moshi/moshi/models/lm.py, lines 77-84) restores these tensors using torch.load(), moving them directly to the target device.
Automatic File Format Detection
The PersonaPlex server (moshi/moshi/server.py, lines 164-168) automatically routes requests based on file extension:
if voice_prompt_path.endswith('.pt'):
self.lm_gen.load_voice_prompt_embeddings(voice_prompt_path)
else:
self.lm_gen.load_voice_prompt(voice_prompt_path)
This logic ensures .pt files trigger the embedding loader, while any other extension routes through the audio pipeline.
Saving Voice Embeddings for Reuse
To generate reusable .pt checkpoints from raw audio, enable the --save-voice-prompt-embeddings flag. The saving mechanism (located in moshi/moshi/models/lm.py, lines 55-61) serializes both the stacked embeddings and streaming cache:
torch.save(
{
"embeddings": torch.stack(saved_embeddings, dim=0).detach().cpu(),
"cache": self._streaming_state.cache,
},
splitext(self.voice_prompt)[0] + ".pt",
)
Practical Implementation Examples
Loading Raw Audio Voice Prompts
Import LMGen and process a WAV file through the audio pipeline:
from moshi.models.lm import LMGen
lm = LMGen(...)
lm.load_voice_prompt("assets/voice_prompts/speaker_voice.wav")
The audio undergoes resampling and normalization via sphn.read before voice encoding occurs.
Loading Pre-Computed Embedding Checkpoints
For production deployments requiring minimal latency, load cached embeddings directly:
from moshi.models.lm import LMGen
lm = LMGen(...)
lm.load_voice_prompt_embeddings("assets/voice_prompts/speaker_voice.pt")
This bypasses audio preprocessing and immediately applies the stored voice characteristics.
Command-Line Interface Usage
Specify either format using the --voice-prompt argument in the offline inference script:
Pre-computed embeddings:
python -m moshi.offline \
--voice-prompt "speaker.pt" \
--input-wav "input.wav" \
--output-wav "output.wav"
Raw audio processing:
python -m moshi.offline \
--voice-prompt "speaker.wav" \
--input-wav "input.wav" \
--output-wav "output.wav"
Summary
- Raw audio files (WAV, FLAC, MP3) supported via
sphn.readare processed on-the-fly throughLMGen.load_voice_prompt()inmoshi/moshi/models/lm.py - PyTorch checkpoint files (
.pt) containing"embeddings"and"cache"tensors load instantly viaLMGen.load_voice_prompt_embeddings() - The server automatically detects format by checking for the
.ptextension inmoshi/moshi/server.py - Embedding checkpoints are generated using
torch.save()with the--save-voice-prompt-embeddingsflag for reuse across sessions
Frequently Asked Questions
Can I use MP3 files directly with PersonaPlex without converting to WAV?
Yes. PersonaPlex accepts any audio format that the sphn.read library supports, which includes MP3, FLAC, and WAV. The LMGen.load_voice_prompt() method handles format detection, resampling, and normalization automatically when loading raw audio inputs.
What is the performance difference between raw audio and .pt checkpoint files?
Pre-computed .pt checkpoints eliminate runtime audio encoding latency because the embeddings are loaded directly via torch.load(). Raw audio requires processing through the full encode pipeline including resampling in load_audio(), making .pt files significantly faster for repeated inference with the same voice.
How do I create a reusable .pt voice embedding checkpoint?
Enable the --save-voice-prompt-embeddings flag when running inference with a raw audio file. The system saves a checkpoint containing the processed embeddings and streaming cache to a .pt file adjacent to your input audio, which can be loaded later via load_voice_prompt_embeddings().
What data is stored inside a PersonaPlex .pt checkpoint file?
Each checkpoint contains two PyTorch tensors: "embeddings" (a stacked tensor of voice representations) and "cache" (the associated streaming state). These are serialized using torch.save() in moshi/moshi/models/lm.py and restored exactly to the model's device when loaded.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →