How the LRU Cache Optimizes Voice Prompts in Pocket-TTS with `_cached_get_state_for_audio_prompt()`
The _cached_get_state_for_audio_prompt() method uses Python's @lru_cache(maxsize=2) decorator to memorize the last two voice prompt states, avoiding redundant downloads, audio encoding, and Flow-LM computations when reusing the same voice reference.
When cloning voices with the TTSModel class in kyutai-labs/pocket-tts, extracting a model state from an audio prompt involves computationally expensive steps: downloading the audio (if a URL is supplied), resampling, encoding with the Mimi codec, and running a forward pass through the Flow-LM. The _cached_get_state_for_audio_prompt() method eliminates this overhead for repeated prompts by caching the resulting dictionary containing hidden states and positional information.
How _cached_get_state_for_audio_prompt() Implements LRU Caching
In pocket_tts/models/tts_model.py, the class defines a private method that wraps the public get_state_for_audio_prompt() API:
@lru_cache(maxsize=2)
def _cached_get_state_for_audio_prompt(
self, audio_conditioning: Path | str | torch.Tensor, truncate: bool = False
) -> dict:
return self.get_state_for_audio_prompt(audio_conditioning, truncate)
This decorator intercepts calls and stores return values in a hash table keyed by the method arguments.
Cache Key Construction
The LRU cache generates keys from the method's call signature. Because audio_conditioning accepts Path, str (URLs or local paths), or torch.Tensor objects, these arguments must be hashable. While Path and str are natively hashable, torch.Tensor support ensures that even tensor-based conditioning can serve as a cache key, though strings and paths represent the typical use case. When identical arguments are passed, the cache returns the previously computed dictionary instantly without repeating the I/O or model inference.
LRU Eviction Policy with maxsize=2
The maxsize=2 parameter restricts the cache to two distinct entries. When a third unique prompt is requested, Python's LRU mechanism evicts the least-recently-used entry to maintain strict memory bounds. This design optimizes for the common scenario where users alternate between two voice personas or repeatedly refine generation with the same prompt, without allowing unbounded memory growth.
Per-Instance Cache Isolation
Because the @lru_cache decorator is applied to a bound method, each TTSModel instance maintains its own independent cache. This isolation prevents cross-contamination between different model configurations or loaded checkpoints. According to the source code in pocket_tts/models/tts_model.py, the cache lives on the instance method, ensuring that loading a new model or creating multiple TTSModel objects does not share stale state between them.
Integration with the Public API and CLI
The public API remains clean through get_state_for_audio_prompt(), while internal performance optimizations use the cached variant. In pocket_tts/main.py, the CLI entry point leverages the cached method directly when processing the --voice argument:
model_state = tts_model._cached_get_state_for_audio_prompt(voice_url)
This implementation detail allows command-line users to benefit from instant voice state retrieval on repeated generations without modifying their interaction pattern.
Complete Working Example
The following pattern demonstrates cache hits and misses in practice:
from pocket_tts import TTSModel
# Load the model once
model = TTSModel.load_model()
# First call – heavy processing (download, encode, run Flow-LM)
voice_state_a = model._cached_get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Subsequent call with identical arguments – cache hit, near-instant return
voice_state_b = model._cached_get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Different prompt – cache miss, triggers full processing
voice_state_c = model._cached_get_state_for_audio_prompt(
"./my_other_voice.wav"
)
# Third distinct prompt evicts the first entry due to maxsize=2
voice_state_d = model._cached_get_state_for_audio_prompt(
"./third_voice.wav"
) # The first prompt would now require recomputation if requested again
Summary
_cached_get_state_for_audio_prompt()wraps the expensiveget_state_for_audio_prompt()method with@lru_cache(maxsize=2)located inpocket_tts/models/tts_model.py.- The cache stores only two entries to balance performance gains against memory constraints, automatically evicting the least-recently-used state when a third distinct prompt arrives.
- Cache isolation occurs at the instance level, preventing state leakage between different
TTSModelobjects. - The CLI implementation in
pocket_tts/main.pycalls the cached method directly to accelerate repeated voice generations from the same prompt. - Hashable arguments (
Path,str, ortorch.Tensor) serve as cache keys, enabling instant retrieval of dictionaries containing pre-computed hidden states and positional embeddings.
Frequently Asked Questions
Why is the cache size limited to only two entries?
The maxsize=2 restriction prevents unbounded memory growth while optimizing for the most common use case: alternating between two voice personas or retrying generation with the same prompt. As implemented in kyutai-labs/pocket-tts, this limit ensures that memory consumption remains predictable even during long-running inference sessions.
What types of audio conditioning arguments work with the cache?
The method accepts Path objects, str (URLs or local file paths), or torch.Tensor objects. Because Python's functools.lru_cache requires hashable keys, Path and str arguments provide natural caching behavior, while torch.Tensor objects are supported for advanced use cases. Non-hashable types would raise a TypeError if passed directly.
Is the cache shared between multiple TTSModel instances?
No. Each TTSModel instance maintains its own private cache because the @lru_cache decorator is applied to the bound method. Creating multiple model instances or loading different checkpoints results in completely separate cache namespaces, preventing cross-talk between models.
Does the LRU cache persist across Python sessions?
No, the cache exists only in memory for the lifetime of the TTSModel instance. When the Python process terminates or the instance is garbage collected, the cached voice states are lost. Subsequent runs must recompute the initial prompt states, though they will again benefit from caching within that session.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →