How Voice Cloning Works in pocket‑tts Using `get_state_for_audio_prompt()`

Voice cloning in pocket‑tts works by extracting a compact VoiceState from a reference audio file via get_state_for_audio_prompt(), which encodes the audio through the Mimi neural codec and caches the result for efficient reuse during streaming generation.

The pocket‑tts repository by Kyutai Labs implements zero-shot voice cloning through a conditioning mechanism that leverages the Mimi audio codec and a Flow-based Language Model (Flow-LM). At the heart of this system is the get_state_for_audio_prompt() method defined in pocket_tts/models/tts_model.py, which transforms raw audio into a structured voice state that persists throughout the text-to-speech generation process.

The Voice Cloning Pipeline

The get_state_for_audio_prompt() method orchestrates a five-stage pipeline to convert a reference waveform into a conditioning signal. Each stage is implemented across specific modules in the codebase.

Step 1: Audio Loading and Preprocessing

The process begins in pocket_tts/data/audio.py, where the input waveform undergoes standardization before neural encoding. The system reads the raw audio file, resamples it to the model's internal 24 kHz sample rate, and applies normalization to ensure consistent amplitude levels.

  • audio.read_wav: Loads the raw bytes from the file path.
  • audio.resample: Converts the waveform to 24 kHz to match the Mimi encoder's expectations.

Step 2: Mimi Codec Encoding

The preprocessed audio is fed into the Mimi encoder located in pocket_tts/models/mimi.py. The MimiModel.encode() function performs a dual extraction:

  • latent_codes: A sequence of discrete tokens capturing acoustic content and temporal structure.
  • speaker_embedding: A continuous vector representing the speaker's unique timbre and vocal characteristics.

This compression is critical because it distills the reference audio into a compact representation that the generative model can efficiently process.

Step 3: VoiceState Construction

The extracted latent_codes and speaker_embedding are packaged into a VoiceState dataclass defined within tts_model.py. This immutable structure serves as the standardized interface between the cloning frontend and the generation backend, ensuring that all acoustic conditioning information travels together through the pipeline.

Step 4: LRU Caching for Performance

Voice cloning is computationally expensive due to the Mimi encoder's forward pass. To optimize repeated use of the same reference voice, the method is wrapped with @functools.lru_cache(maxsize=64) on an internal helper called _cached_get_state_for_audio_prompt. When you call get_state_for_audio_prompt() with the same file path, the system returns the cached VoiceState instantly, eliminating redundant encoding overhead.

Step 5: Conditioning the Flow-LM

During generation, generate_audio_stream() accepts the VoiceState object via the voice_state parameter. The method injects these embeddings into the Flow-LM's internal state at every streaming step, ensuring the synthesized audio maintains consistent speaker identity with the reference prompt throughout the entire utterance.

Practical Implementation with get_state_for_audio_prompt()

The public API abstracts the complexity of the pipeline behind a single method call. Below are practical patterns for implementing voice cloning in your applications.

Basic Voice Cloning

This example demonstrates the complete workflow from loading a reference voice to generating cloned speech:

from pocket_tts import TTSModel

# Initialize the model (downloads weights on first run)

tts = TTSModel()

# Path to your reference audio (WAV format, any length)

voice_path = "examples/voice_prompt.wav"

# Extract the voice state - this performs the cloning

voice_state = tts.get_state_for_audio_prompt(voice_path)

# Generate speech using the cloned voice identity

text = "Hello, this is a cloned voice speaking."
audio_chunks = tts.generate_audio_stream(
    text,
    voice_state=voice_state,  # Condition on the cloned voice

    temperature=0.7,
)

# Save the streamed output

with open("output.wav", "wb") as f:
    for chunk in audio_chunks:
        f.write(chunk)

Reusing Cached Voice States

Because get_state_for_audio_prompt() implements LRU caching, subsequent calls with the same path return instantly:

from pocket_tts import TTSModel

tts = TTSModel()
voice_path = "examples/voice_prompt.wav"

# First call triggers Mimi encoding (slow)

base_voice = tts.get_state_for_audio_prompt(voice_path)

# Subsequent calls hit the cache (instant)

for sentence in ["First sentence.", "Second sentence.", "Third sentence."]:
    audio = tts.generate_audio_stream(
        sentence,
        voice_state=tts.get_state_for_audio_prompt(voice_path),
    )
    # Process audio chunks...

Key Source Files and Architecture

Understanding the repository structure helps when customizing or debugging the voice cloning pipeline:

Summary

  • get_state_for_audio_prompt() serves as the entry point for voice cloning, converting reference audio into a reusable VoiceState object.
  • The method relies on the Mimi codec (mimi.py) to extract discrete latent tokens and continuous speaker embeddings from the input waveform.
  • LRU caching (maxsize=64) prevents redundant encoding of identical voice prompts, significantly improving performance for batch generation.
  • The resulting VoiceState conditions the Flow-LM during generate_audio_stream(), ensuring consistent speaker characteristics throughout synthesis.
  • Audio preprocessing happens in audio.py, which handles resampling to 24 kHz and normalization.

Frequently Asked Questions

What is the purpose of get_state_for_audio_prompt()?

The method extracts a structured representation of a reference voice that the TTS model can use as a conditioning signal. It encapsulates the entire preprocessing and encoding pipeline, returning a VoiceState object that contains both the acoustic content (latent codes) and speaker identity (embedding) needed to guide the Flow-LM during generation.

How does the Mimi codec contribute to voice cloning?

The Mimi codec (pocket_tts/models/mimi.py) functions as a neural audio compressor. Its encoder produces two distinct outputs: discrete latent codes that capture the temporal acoustic structure, and a continuous speaker embedding that encodes timbre. Together, these allow the generative model to reproduce both the speaking style and the specific vocal characteristics of the reference audio.

Why is the voice state cached with lru_cache?

Encoding a voice prompt requires a full forward pass through the Mimi neural network, which involves significant computation. By wrapping the internal _cached_get_state_for_audio_prompt() with @functools.lru_cache(maxsize=64), the system stores the resulting VoiceState in memory. Subsequent requests for the same file path return the cached object instantly, reducing latency when generating multiple utterances with the same cloned voice.

Can I use multiple different voice prompts in the same session?

Yes. You can call get_state_for_audio_prompt() with different file paths to create distinct VoiceState objects, and the cache will retain them up to the 64-entry limit. Pass the appropriate voice_state to each generate_audio_stream() call to switch between different cloned voices or the default speaker at runtime.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →