# How to Export and Reload Voice States Using Safetensors for Fast Voice Loading in Pocket-TTS

> Export and reload Voice States with safetensors in Pocket-TTS for instant, fast voice cloning. Bypass encoding with export_model_state and get_state_for_audio_prompt.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-11

---

**Pocket-TTS allows you to export voice states to `.safetensors` files using `export_model_state` and reload them instantly via `get_state_for_audio_prompt`, bypassing costly audio encoding for fast voice cloning.**

When working with the `kyutai-labs/pocket-tts` repository, extracting a speaker's voice characteristics from raw audio requires expensive processing through the Mimi codec and flow-LM transformer. To eliminate this overhead, you can export and reload voice states using safetensors, enabling millisecond-fast voice cloning after the initial extraction.

## Understanding Voice States in Pocket-TTS

A **voice state** (also called a model state) is a nested dictionary of tensors containing the hidden KV caches and positional offsets required by the flow-LM and Mimi modules. When you call `get_state_for_audio_prompt` on an audio file, the system resamples the audio, encodes it through the Mimi codec, projects it into the flow-LM latent space, and initializes these states using `init_states`.

This initialization process is computationally expensive because it involves a complete forward pass through the flow model. For production deployments or interactive applications, repeating this work for every session creates unacceptable latency.

## Exporting Voice States to Safetensors

### The `export_model_state` Function

The `export_model_state` function, located in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 4747–4752), serializes the nested state dictionary into a flat format compatible with the Safetensors specification. The function iterates over every tensor in the state, constructs flat keys using the pattern `"module_name/tensor_key"`, and writes the complete mapping using `safetensors.torch.save_file`.

This process preserves the hierarchical module structure while creating a binary, zero-copy container that loads efficiently without arbitrary code execution risks.

```python
from pocket_tts import TTSModel, export_model_state

# Load the model (default English config)

model = TTSModel.load_model()

# Create a state from any audio source (local file, HF URL, etc.)

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Persist the state – this file can be reused across runs

export_model_state(voice_state, "alba_casual.safetensors")
print("Voice state exported to alba_casual.safetensors")

```

## Reloading Voice States for Instant Cloning

### Automatic Detection with `_import_model_state`

When you pass a file path ending in `.safetensors` to `get_state_for_audio_prompt`, the method detects the suffix (as implemented in lines 445–452 of [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)) and immediately routes to `_import_model_state` instead of processing audio. The `_import_model_state` helper (lines 4755–4772) opens the file using `safetensors.safe_open`, splits each flat key back into its hierarchical components, and translates tensors onto the requested device.

A special case handles the `offset` tensor—which tracks KV-cache positions—to ensure the generator remains synchronized with the original voice extraction.

```python
from pocket_tts import TTSModel

# Load the same model (or a different one with compatible config)

model = TTSModel.load_model()

# Directly load the previously exported state – no audio processing needed

voice_state = model.get_state_for_audio_prompt("alba_casual.safetensors")

# Use the state to generate speech in the same voice

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Hello, this is a fast‑loaded voice!",
    frames_after_eos=2,
    copy_state=True,
)

```

## Complete Workflow Example

This pattern separates voice extraction from inference, allowing you to pre-process voices in one environment and deploy them in another:

```python

# ---- Process A: Export -------------------------------------------------

from pocket_tts import TTSModel, export_model_state

model_a = TTSModel.load_model()
state_a = model_a.get_state_for_audio_prompt("my_voice.wav")
export_model_state(state_a, "my_voice.safetensors")

# ---- Process B (different script or machine) -------------------------

from pocket_tts import TTSModel

model_b = TTSModel.load_model()
state_b = model_b.get_state_for_audio_prompt("my_voice.safetensors")
audio_b = model_b.generate_audio(state_b, "Quick reload test")

```

The `get_state_for_audio_prompt` method utilizes `download_if_necessary` from [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py) to support remote URLs for both audio files and pre-exported safetensors, making deployment across distributed systems seamless.

## Performance Benefits and Technical Details

The **Safetensors** format enables memory-mapped loading that is CPU-only and avoids the GPU inference overhead required during initial voice extraction. Because the exported state contains pre-computed KV caches from [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py) and encoded representations from [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py), reloading eliminates resampling, encoding, and forward-pass computations entirely.

This architecture decouples the expensive voice-cloning preparation from real-time synthesis, making it practical to maintain large libraries of speaker voices that initialize instantly regardless of original audio length.

## Summary

- **Voice states** are dictionaries of tensors containing KV caches and positional offsets from the Mimi codec and flow-LM transformer.
- Use `export_model_state` in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 4747–4752) to serialize states to `.safetensors` files.
- Pass `.safetensors` paths directly to `get_state_for_audio_prompt` to trigger `_import_model_state` (lines 4755–4772) and skip audio processing.
- The Safetensors format provides fast, safe, zero-copy reloading suitable for production voice cloning pipelines.

## Frequently Asked Questions

### What file format does Pocket-TTS use for voice states?

Pocket-TTS uses the **Safetensors** format (`.safetensors`), a binary container that stores tensors in a memory-mapped, zero-copy structure. This format prevents arbitrary code execution during loading and enables fast CPU-only deserialization, as implemented in `export_model_state` using `safetensors.torch.save_file`.

### Can I share exported voice states between different machines?

Yes. Exported `.safetensors` files are portable and contain no machine-specific metadata. You can generate a voice state on one system using `export_model_state`, transfer the file, and load it on another machine using `get_state_for_audio_prompt` with the same model configuration, bypassing the need to transfer original audio files.

### Why is reloading from safetensors faster than processing audio?

Reloading from safetensors skips the computationally expensive steps of audio resampling, Mimi encoding (defined in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)), and flow-LM forward passes (implemented in [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py)). The `_import_model_state` function simply maps pre-computed tensors into memory, making voice initialization instantaneous regardless of the original audio duration.

### Do I need the original audio file after exporting the state?

No. Once you have exported the voice state to a `.safetensors` file, the original audio file is no longer required for synthesis. The exported state contains all necessary acoustic characteristics, including the hidden KV caches and offset tensors needed to maintain voice consistency during generation.