How to Export and Reload Voice States Using Safetensors for Fast Voice Loading in Pocket-TTS

Pocket-TTS allows you to export voice states to .safetensors files using export_model_state and reload them instantly via get_state_for_audio_prompt, bypassing costly audio encoding for fast voice cloning.

When working with the kyutai-labs/pocket-tts repository, extracting a speaker's voice characteristics from raw audio requires expensive processing through the Mimi codec and flow-LM transformer. To eliminate this overhead, you can export and reload voice states using safetensors, enabling millisecond-fast voice cloning after the initial extraction.

Understanding Voice States in Pocket-TTS

A voice state (also called a model state) is a nested dictionary of tensors containing the hidden KV caches and positional offsets required by the flow-LM and Mimi modules. When you call get_state_for_audio_prompt on an audio file, the system resamples the audio, encodes it through the Mimi codec, projects it into the flow-LM latent space, and initializes these states using init_states.

This initialization process is computationally expensive because it involves a complete forward pass through the flow model. For production deployments or interactive applications, repeating this work for every session creates unacceptable latency.

Exporting Voice States to Safetensors

The export_model_state Function

The export_model_state function, located in pocket_tts/models/tts_model.py (lines 4747–4752), serializes the nested state dictionary into a flat format compatible with the Safetensors specification. The function iterates over every tensor in the state, constructs flat keys using the pattern "module_name/tensor_key", and writes the complete mapping using safetensors.torch.save_file.

This process preserves the hierarchical module structure while creating a binary, zero-copy container that loads efficiently without arbitrary code execution risks.

from pocket_tts import TTSModel, export_model_state

# Load the model (default English config)

model = TTSModel.load_model()

# Create a state from any audio source (local file, HF URL, etc.)

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Persist the state – this file can be reused across runs

export_model_state(voice_state, "alba_casual.safetensors")
print("Voice state exported to alba_casual.safetensors")

Reloading Voice States for Instant Cloning

Automatic Detection with _import_model_state

When you pass a file path ending in .safetensors to get_state_for_audio_prompt, the method detects the suffix (as implemented in lines 445–452 of pocket_tts/models/tts_model.py) and immediately routes to _import_model_state instead of processing audio. The _import_model_state helper (lines 4755–4772) opens the file using safetensors.safe_open, splits each flat key back into its hierarchical components, and translates tensors onto the requested device.

A special case handles the offset tensor—which tracks KV-cache positions—to ensure the generator remains synchronized with the original voice extraction.

from pocket_tts import TTSModel

# Load the same model (or a different one with compatible config)

model = TTSModel.load_model()

# Directly load the previously exported state – no audio processing needed

voice_state = model.get_state_for_audio_prompt("alba_casual.safetensors")

# Use the state to generate speech in the same voice

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Hello, this is a fast‑loaded voice!",
    frames_after_eos=2,
    copy_state=True,
)

Complete Workflow Example

This pattern separates voice extraction from inference, allowing you to pre-process voices in one environment and deploy them in another:


# ---- Process A: Export -------------------------------------------------

from pocket_tts import TTSModel, export_model_state

model_a = TTSModel.load_model()
state_a = model_a.get_state_for_audio_prompt("my_voice.wav")
export_model_state(state_a, "my_voice.safetensors")

# ---- Process B (different script or machine) -------------------------

from pocket_tts import TTSModel

model_b = TTSModel.load_model()
state_b = model_b.get_state_for_audio_prompt("my_voice.safetensors")
audio_b = model_b.generate_audio(state_b, "Quick reload test")

The get_state_for_audio_prompt method utilizes download_if_necessary from pocket_tts/utils/utils.py to support remote URLs for both audio files and pre-exported safetensors, making deployment across distributed systems seamless.

Performance Benefits and Technical Details

The Safetensors format enables memory-mapped loading that is CPU-only and avoids the GPU inference overhead required during initial voice extraction. Because the exported state contains pre-computed KV caches from pocket_tts/models/flow_lm.py and encoded representations from pocket_tts/models/mimi.py, reloading eliminates resampling, encoding, and forward-pass computations entirely.

This architecture decouples the expensive voice-cloning preparation from real-time synthesis, making it practical to maintain large libraries of speaker voices that initialize instantly regardless of original audio length.

Summary

  • Voice states are dictionaries of tensors containing KV caches and positional offsets from the Mimi codec and flow-LM transformer.
  • Use export_model_state in pocket_tts/models/tts_model.py (lines 4747–4752) to serialize states to .safetensors files.
  • Pass .safetensors paths directly to get_state_for_audio_prompt to trigger _import_model_state (lines 4755–4772) and skip audio processing.
  • The Safetensors format provides fast, safe, zero-copy reloading suitable for production voice cloning pipelines.

Frequently Asked Questions

What file format does Pocket-TTS use for voice states?

Pocket-TTS uses the Safetensors format (.safetensors), a binary container that stores tensors in a memory-mapped, zero-copy structure. This format prevents arbitrary code execution during loading and enables fast CPU-only deserialization, as implemented in export_model_state using safetensors.torch.save_file.

Can I share exported voice states between different machines?

Yes. Exported .safetensors files are portable and contain no machine-specific metadata. You can generate a voice state on one system using export_model_state, transfer the file, and load it on another machine using get_state_for_audio_prompt with the same model configuration, bypassing the need to transfer original audio files.

Why is reloading from safetensors faster than processing audio?

Reloading from safetensors skips the computationally expensive steps of audio resampling, Mimi encoding (defined in pocket_tts/models/mimi.py), and flow-LM forward passes (implemented in pocket_tts/models/flow_lm.py). The _import_model_state function simply maps pre-computed tensors into memory, making voice initialization instantaneous regardless of the original audio duration.

Do I need the original audio file after exporting the state?

No. Once you have exported the voice state to a .safetensors file, the original audio file is no longer required for synthesis. The exported state contains all necessary acoustic characteristics, including the hidden KV caches and offset tensors needed to maintain voice consistency during generation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →