How Pocket-TTS Uses an LRU Cache to Manage Voice Prompt State

Pocket-TTS caches expensive voice-prompt encoding operations using Python's functools.lru_cache decorator on a private helper method, storing up to two recent model states to eliminate redundant audio processing and KV-cache initialization.

When working with the kyutai-labs/pocket-tts library, converting a speaker's audio sample into a usable generation state requires significant computation. The library employs an LRU cache to manage voice prompt state, ensuring that repeated use of the same voice sample avoids redundant encoding overhead while maintaining strict memory boundaries.

Why Voice-Prompt Encoding Requires Caching

Generating speech with a specific voice requires encoding an audio prompt into a model state before inference can begin. This pipeline involves loading the audio file, resampling, passing it through the Mimi codec, projecting into the FlowLM latent space, and building a KV-cache for the transformer. Because this sequence is computationally expensive, Pocket-TTS memoizes the results to accelerate subsequent generations with the same voice.

Implementation of the LRU Cache in pocket_tts/models/tts_model.py

The caching mechanism is implemented in the core model class using Python’s standard library.

The Cached Wrapper Method

At lines 81–86 of pocket_tts/models/tts_model.py, the library defines a private wrapper method decorated with functools.lru_cache:

@lru_cache(maxsize=2)
def _cached_get_state_for_audio_prompt(self, audio_conditioning, truncate=False):
    return self.get_state_for_audio_prompt(audio_conditioning, truncate)

This wrapper intercepts calls to the expensive state-generation logic. When an application requests a voice prompt state, the public get_state_for_audio_prompt method forwards the request to this cached helper. If the arguments match a recent entry, the cached state dictionary returns immediately without re-executing the encoding pipeline.

Cache Keys and Hashability

The cache key derives directly from the method arguments:

  • audio_conditioning – A Path, URL string, or tensor identifying the prompt audio
  • truncate – A boolean flag controlling audio truncation

Because these arguments are hashable, each distinct combination generates a unique cache entry.

Cache Contents and Structure

When the cache misses, the underlying get_state_for_audio_prompt method (implementation begins around line 891) constructs a comprehensive state dictionary containing:

  • Encoder-projected audio conditioning tensors – The processed audio embeddings ready for the transformer
  • Initialized KV-cache – Pre-computed key-value caches for the FlowLM transformer (init_states)
  • Position offsets – Tracking indices for autoregressive generation
  • Speaker-projection weights – Voice-specific projection parameters

This dictionary represents the complete initialization state required to continue generation with the specified voice.

Cache Limits and Eviction Policy

The @lru_cache(maxsize=2) configuration imposes a hard limit of two entries:

  • When a third distinct prompt is requested, the least-recently-used entry evicts automatically
  • This size accommodates common use cases involving one primary voice or a fallback speaker
  • The small bound prevents unbounded memory growth while delivering significant latency reductions for repeated calls

Practical Usage Example

The following demonstrates the cache behavior in practice:

from pocket_tts import TTSModel

# Load the model once (weights are cached on disk)

model = TTSModel.load_model()

# First use – expensive encoding occurs

voice_state_a = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Second use of the same prompt – served from LRU cache

voice_state_a2 = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

assert voice_state_a is voice_state_a2  # Identical cached object

# Third distinct prompt evicts the least-recently-used entry

voice_state_b = model.get_state_for_audio_prompt("./my_other_voice.wav")

Thread Safety Considerations

The LRU cache lives on the TTSModel instance. Because Pocket-TTS generation is not thread-safe by design, the cache requires no additional locking mechanisms. Concurrent accesses are avoided by using separate model instances per thread rather than sharing a single instance across threads.

Summary

  • Pocket-TTS uses functools.lru_cache on _cached_get_state_for_audio_prompt in pocket_tts/models/tts_model.py to memoize voice-prompt states
  • The cache stores complete model state dictionaries containing KV-caches, audio projections, and position offsets generated by get_state_for_audio_prompt
  • Maximum capacity is strictly limited to two entries (maxsize=2), automatically evicting the least-recently-used prompt when exceeded
  • Cache keys consist of the hashable audio_conditioning identifier and truncate boolean flag
  • The implementation assumes single-threaded access per model instance; concurrent generation requires separate TTSModel instances

Frequently Asked Questions

What is the maximum number of voice prompts Pocket-TTS caches simultaneously?

The LRU cache stores exactly two voice prompt states. The @lru_cache(maxsize=2) decorator at lines 81–86 of pocket_tts/models/tts_model.py enforces this limit, automatically discarding the least-recently-used entry when a third distinct prompt is requested.

Which method in Pocket-TTS is responsible for caching voice prompt states?

The private method _cached_get_state_for_audio_prompt handles caching. This wrapper, defined in pocket_tts/models/tts_model.py, delegates to the public get_state_for_audio_prompt method on cache misses while returning memoized results for recent prompts.

Does the Pocket-TTS LRU cache work across multiple threads?

No. The cache resides on the TTSModel instance, and Pocket-TTS generation is not thread-safe by design. For concurrent processing, instantiate separate TTSModel objects per thread rather than sharing cached states across threads.

What components are included in the cached voice prompt state?

The cached dictionary contains encoder-projected audio conditioning tensors, initialized KV-caches for the FlowLM transformer (init_states), generation position offsets, and speaker-projection weights. This state is constructed by the get_state_for_audio_prompt implementation beginning at line 891.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →