# Difference Between NAT and VAR Voice Prompt Embeddings in PersonaPlex

> Understand the difference between NAT and VAR voice prompt embeddings in NVIDIA PersonaPlex. Discover how NAT offers natural speech and VAR provides expressive diversity for your AI personas.

- Repository: [NVIDIA Corporation/personaplex](https://github.com/NVIDIA/personaplex)
- Tags: deep-dive
- Published: 2026-04-07

---

**NAT embeddings provide natural, conversational speech patterns while VAR embeddings deliver more expressive, stylistically diverse vocal characteristics, though both use identical PyTorch checkpoint formats and loading mechanisms in the PersonaPlex inference pipeline.**

The NVIDIA PersonaPlex repository implements two distinct families of pre-computed voice-prompt embeddings that condition the Moshi 7B language model's audio generation. Understanding the difference between NAT and VAR voice prompt embeddings helps you select the right persona for your application, whether you need stable, natural-sounding speech or more distinctive, character-rich vocal outputs.

## What Are NAT and VAR Voice Families?

According to the [`README.md`](https://github.com/NVIDIA/personaplex/blob/main/README.md) in the repository root, PersonaPlex organizes voice prompts into two categories based on their acoustic training distributions.

### NAT (Natural) Embeddings

**NAT** (Natural) embeddings use identifier prefixes `NATF*` for female voices and `NATM*` for male voices. These embeddings derive from high-quality recordings selected to sound neutral, smooth, and "human-like," matching the conversational style of the original Moshi 7B base model. The pre-computed tensors capture stable, predictable speech patterns ideal for assistant applications requiring consistent, natural-sounding output.

### VAR (Variety) Embeddings

**VAR** (Variety) embeddings use prefixes `VARF*` and `VARM*` and originate from a broader, intentionally varied dataset. These embeddings encode richer acoustic patterns including diverse prosody, pitch contours, and speaker characteristics. When loaded, VAR embeddings produce more distinctive, sometimes "character-ful" voices suitable for applications requiring expressive or stylistically diverse speech outputs.

Both families ship as serialized PyTorch checkpoints (`.pt` files) containing identical internal structures.

## Technical Implementation and File Structure

Regardless of family, every voice prompt checkpoint stores the same dictionary structure. According to the implementation in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py), each `.pt` file contains:

```python
{
    "embeddings": <Tensor of voice-prompt frame embeddings>,
    "cache":      <Tensor representing the LM's internal cache state>
}

```

The `load_voice_prompt_embeddings` method at lines 77-84 of [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py) deserializes these checkpoints directly into the model's inference state:

```python
def load_voice_prompt_embeddings(self, path: str):
    self.voice_prompt = path
    state = torch.load(path)
    self.voice_prompt_audio = None
    self.voice_prompt_embeddings = state["embeddings"].to(self.lm_model.device)
    self.voice_prompt_cache = state["cache"].to(self.lm_model.device)

```

This loading mechanism bypasses the costly on-the-fly audio encoding pipeline, making `.pt` files the preferred production path for both NAT and VAR families.

## Server-Side Loading Logic

The PersonaPlex server determines whether to load pre-computed embeddings or encode raw audio dynamically. In [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py) (lines 48-66), the request handling logic checks the file extension:

```python
if self.lm_gen.voice_prompt != voice_prompt_path:
    if voice_prompt_path.endswith('.pt'):
        self.lm_gen.load_voice_prompt_embeddings(voice_prompt_path)
    else:
        self.lm_gen.load_voice_prompt(voice_prompt_path)

```

When the query parameter `voice_prompt` specifies a `.pt` file (such as `NATF0.pt` or `VARM3.pt`), the server invokes `load_voice_prompt_embeddings`. For raw audio files (e.g., `.wav`), it falls back to `load_voice_prompt`, which extracts embeddings online. Both NAT and VAR families use the fast `.pt` path exclusively.

## Practical Usage Examples

You can specify either family via command-line arguments or the web interface.

### Running Offline Inference with NAT

To generate speech using a natural female voice prompt:

```bash
HF_TOKEN=$HF_TOKEN \
python -m moshi.offline \
  --voice-prompt "NATF2.pt" \
  --input-wav "assets/test/input_assistant.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

```

### Running Offline Inference with VAR

To use a varied male voice with more character:

```bash
HF_TOKEN=$HF_TOKEN \
python -m moshi.offline \
  --voice-prompt "VARM3.pt" \
  --input-wav "assets/test/input_service.wav" \
  --seed 42424242 \
  --output-wav "output.wav" \
  --output-text "output.json"

```

### Selecting Voices in the Web UI

The React frontend exposes all 16 available prompts (8 NAT and 8 VAR) through a dropdown component. In [`client/src/pages/Conversation/components/ModelParams/ModelParams.tsx`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Conversation/components/ModelParams/ModelParams.tsx), the selection maps directly to the filename:

```tsx
<select
  value={modelParams.voicePrompt}
  onChange={e => setModelParams({ ...modelParams, voicePrompt: e.target.value })}
>
  <option value="NATF0.pt">NATF0.pt</option>
  <option value="NATF1.pt">NATF1.pt</option>
  <option value="VARF0.pt">VARF0.pt</option>
  <option value="VARM4.pt">VARM4.pt</option>
</select>

```

The complete list of available prompts is hard-coded in [`client/src/pages/Queue/Queue.tsx`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Queue/Queue.tsx) for quick reference.

## Why Two Families Exist?

PersonaPlex maintains separate NAT and VAR families for two primary reasons:

1. **Training Data Distribution**: NAT embeddings source from recordings targeting neutral, natural conversation, aligning with the base model's training distribution. VAR embeddings deliberately sample from diverse acoustic styles to demonstrate the model's ability to adopt distinct personas.

2. **User-Facing Selection**: The web UI groups prompts by family, allowing users to choose between stable, predictable voices (NAT) or more expressive, varied characteristics (VAR) without changing the underlying technical implementation.

Both families remain interchangeable at inference time; the sole difference lies in the acoustic style encoded by the embedding tensors.

## Summary

- **NAT (Natural)** embeddings (`NATF*`/`NATM*`) provide stable, neutral, conversational speech patterns derived from high-quality natural recordings.
- **VAR (Variety)** embeddings (`VARF*`/`VARM*`) offer expressive, diverse vocal characteristics sourced from varied acoustic datasets.
- Both families use identical `.pt` checkpoint formats containing `embeddings` and `cache` tensors, loaded via `load_voice_prompt_embeddings` in [`moshi/models/lm.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/models/lm.py).
- The server at [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py) routes `.pt` files to the fast embedding path, while raw audio undergoes online encoding.
- Select either family via command-line flags or the React dropdown in [`client/src/pages/Conversation/components/ModelParams/ModelParams.tsx`](https://github.com/NVIDIA/personaplex/blob/main/client/src/pages/Conversation/components/ModelParams/ModelParams.tsx).

## Frequently Asked Questions

### Can I switch between NAT and VAR embeddings during the same session?

Yes. Both families use the same checkpoint format and loading mechanism. You can change the `voice_prompt` parameter between requests to switch from a NAT to a VAR voice, and the server will load the new embeddings via `load_voice_prompt_embeddings` without requiring model restarts.

### Do NAT embeddings only support female voices and VAR only male voices?

No. Both families support gender variants. NAT uses prefixes `NATF*` (female) and `NATM*` (male), while VAR uses `VARF*` (female) and `VARM*` (male). Each family contains four female and four male voice options in the standard distribution documented in [`README.md`](https://github.com/NVIDIA/personaplex/blob/main/README.md).

### What is the performance impact of using pre-computed .pt files versus raw audio?

Pre-computed `.pt` files significantly reduce latency. When loading via `load_voice_prompt_embeddings`, the system bypasses the audio encoder entirely, reading tensors directly from disk. Raw audio files trigger `load_voice_prompt`, which processes the waveform through the encoder network on every request, increasing inference time.

### How do I add custom voice prompts to the available NAT or VAR sets?

You can generate custom embeddings by capturing the output of the audio encoder and internal cache state, then serializing them to a `.pt` file containing the `embeddings` and `cache` keys. Place these files in the voice prompt directory configured in [`moshi/server.py`](https://github.com/NVIDIA/personaplex/blob/main/moshi/server.py), and they will be available for selection alongside the standard NAT and VAR families.