# How Pocket-TTS Uses an LRU Cache to Manage Voice Prompt State

> Discover how Pocket-TTS optimizes voice prompt state management with an LRU cache. Learn how it slashes redundant audio processing and KV-cache initialization for faster performance.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: internals
- Published: 2026-07-09

---

**Pocket-TTS caches expensive voice-prompt encoding operations using Python's `functools.lru_cache` decorator on a private helper method, storing up to two recent model states to eliminate redundant audio processing and KV-cache initialization.**

When working with the `kyutai-labs/pocket-tts` library, converting a speaker's audio sample into a usable generation state requires significant computation. The library employs an LRU cache to manage voice prompt state, ensuring that repeated use of the same voice sample avoids redundant encoding overhead while maintaining strict memory boundaries.

## Why Voice-Prompt Encoding Requires Caching

Generating speech with a specific voice requires encoding an audio prompt into a *model state* before inference can begin. This pipeline involves loading the audio file, resampling, passing it through the Mimi codec, projecting into the FlowLM latent space, and building a KV-cache for the transformer. Because this sequence is computationally expensive, Pocket-TTS memoizes the results to accelerate subsequent generations with the same voice.

## Implementation of the LRU Cache in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)

The caching mechanism is implemented in the core model class using Python’s standard library.

### The Cached Wrapper Method

At lines 81–86 of [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), the library defines a private wrapper method decorated with `functools.lru_cache`:

```python
@lru_cache(maxsize=2)
def _cached_get_state_for_audio_prompt(self, audio_conditioning, truncate=False):
    return self.get_state_for_audio_prompt(audio_conditioning, truncate)

```

This wrapper intercepts calls to the expensive state-generation logic. When an application requests a voice prompt state, the public `get_state_for_audio_prompt` method forwards the request to this cached helper. If the arguments match a recent entry, the cached state dictionary returns immediately without re-executing the encoding pipeline.

### Cache Keys and Hashability

The cache key derives directly from the method arguments:

- `audio_conditioning` – A `Path`, URL string, or tensor identifying the prompt audio
- `truncate` – A boolean flag controlling audio truncation

Because these arguments are hashable, each distinct combination generates a unique cache entry.

## Cache Contents and Structure

When the cache misses, the underlying `get_state_for_audio_prompt` method (implementation begins around line 891) constructs a comprehensive state dictionary containing:

- **Encoder-projected audio conditioning tensors** – The processed audio embeddings ready for the transformer
- **Initialized KV-cache** – Pre-computed key-value caches for the FlowLM transformer (`init_states`)
- **Position offsets** – Tracking indices for autoregressive generation
- **Speaker-projection weights** – Voice-specific projection parameters

This dictionary represents the complete initialization state required to continue generation with the specified voice.

## Cache Limits and Eviction Policy

The `@lru_cache(maxsize=2)` configuration imposes a hard limit of two entries:

- When a third distinct prompt is requested, the **least-recently-used** entry evicts automatically
- This size accommodates common use cases involving one primary voice or a fallback speaker
- The small bound prevents unbounded memory growth while delivering significant latency reductions for repeated calls

## Practical Usage Example

The following demonstrates the cache behavior in practice:

```python
from pocket_tts import TTSModel

# Load the model once (weights are cached on disk)

model = TTSModel.load_model()

# First use – expensive encoding occurs

voice_state_a = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Second use of the same prompt – served from LRU cache

voice_state_a2 = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

assert voice_state_a is voice_state_a2  # Identical cached object

# Third distinct prompt evicts the least-recently-used entry

voice_state_b = model.get_state_for_audio_prompt("./my_other_voice.wav")

```

## Thread Safety Considerations

The LRU cache lives on the `TTSModel` instance. Because Pocket-TTS generation is **not thread-safe by design**, the cache requires no additional locking mechanisms. Concurrent accesses are avoided by using separate model instances per thread rather than sharing a single instance across threads.

## Summary

- Pocket-TTS uses `functools.lru_cache` on `_cached_get_state_for_audio_prompt` in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) to memoize voice-prompt states
- The cache stores complete model state dictionaries containing KV-caches, audio projections, and position offsets generated by `get_state_for_audio_prompt`
- Maximum capacity is strictly limited to **two entries** (`maxsize=2`), automatically evicting the least-recently-used prompt when exceeded
- Cache keys consist of the hashable `audio_conditioning` identifier and `truncate` boolean flag
- The implementation assumes single-threaded access per model instance; concurrent generation requires separate `TTSModel` instances

## Frequently Asked Questions

### What is the maximum number of voice prompts Pocket-TTS caches simultaneously?

The LRU cache stores exactly **two** voice prompt states. The `@lru_cache(maxsize=2)` decorator at lines 81–86 of [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) enforces this limit, automatically discarding the least-recently-used entry when a third distinct prompt is requested.

### Which method in Pocket-TTS is responsible for caching voice prompt states?

The private method `_cached_get_state_for_audio_prompt` handles caching. This wrapper, defined in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), delegates to the public `get_state_for_audio_prompt` method on cache misses while returning memoized results for recent prompts.

### Does the Pocket-TTS LRU cache work across multiple threads?

No. The cache resides on the `TTSModel` instance, and Pocket-TTS generation is not thread-safe by design. For concurrent processing, instantiate separate `TTSModel` objects per thread rather than sharing cached states across threads.

### What components are included in the cached voice prompt state?

The cached dictionary contains encoder-projected audio conditioning tensors, initialized KV-caches for the FlowLM transformer (`init_states`), generation position offsets, and speaker-projection weights. This state is constructed by the `get_state_for_audio_prompt` implementation beginning at line 891.