# How Pocket-TTS Streaming Architecture Works with StatefulModule and KV Cache Management

> Understand Pocket-TTS streaming architecture with StatefulModule and KV cache. Achieve real-time, CPU-only speech synthesis by preserving attention keys and values, eliminating redundant computation.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: architecture
- Published: 2026-07-11

---

**Pocket-TTS achieves real-time, CPU-only speech synthesis by generating audio frame-by-frame using StatefulModule to preserve attention keys and values across steps, eliminating redundant computation through incremental KV-cache updates.**

Pocket-TTS (kyutai-labs/pocket-tts) implements a frame-level streaming architecture that enables low-latency text-to-speech on consumer hardware. By leveraging **StatefulModule** and **KV cache management**, the system avoids recomputing full attention matrices for every audio frame, drastically reducing memory usage and computational overhead during generation.

## Frame-by-Frame Generation at 12.5 Hz

Pocket-TTS generates speech **frame-by-frame** at approximately 12.5 Hz (every 80 ms). Rather than processing entire sequences in a single forward pass, the model operates incrementally, producing small audio chunks that can be streamed to output devices with minimal latency. This approach requires maintaining context across successive forward calls without reloading the full sequence history.

## StatefulModule: Centralizing State Preservation

The foundation of the streaming architecture is **`StatefulModule`**, a base class defined in [`pocket_tts/modules/stateful_module.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/stateful_module.py). This class centralizes two critical responsibilities: preserving internal tensors across forward calls and managing the KV cache for attention mechanisms.

### The State Dictionary

Each module inherits a **`self._state`** dictionary that persists any internal tensors required for subsequent generation steps. When a transformer block processes a new token or audio frame, it stores computed keys and values in this state container. According to the source code in [`pocket_tts/modules/stateful_module.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/stateful_module.py), this dictionary acts as the persistent memory for the module's streaming context.

### Resetting State Between Utterances

When generating a new utterance or interrupting current synthesis, the **`reset_state()`** method clears all cached tensors. The `TTSModel.generate_audio_stream()` method automatically triggers this reset at utterance boundaries, ensuring that each new speech sequence starts from a clean slate without residual attention states from previous inputs.

## KV Cache Management in Streaming Transformers

The **KV cache** (key-value cache) mechanism lives entirely within the `StatefulModule` state system. Transformer-based components like **`StreamingTransformer`** and **`StreamingMultiheadAttention`** (implemented in [`pocket_tts/modules/transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/transformer.py)) rely on this cache to avoid recomputing attention weights for all previous positions at every step.

### Incremental Cache Updates

The cache follows a three-phase update pattern:

1. **First step**: The cache is empty. The module computes keys and values from the initial input token or frame and stores them in `self._state`.

2. **Subsequent steps**: The module computes only the new query vector, then concatenates it with cached keys and values from previous steps. The updated cache is stored back into `self._state`.

3. **Reset**: When `reset_state()` is called (either manually or automatically at utterance boundaries), the cached tensors are cleared.

This incremental approach means the attention layer only computes queries for the fresh token while reusing cached keys and values for all previous positions, drastically reducing per-step computational cost.

### Shared Cache Across the Pipeline

Because the KV cache is unified across all streaming modules, the entire pipeline—including the text encoder, flow language model, and audio decoder—can operate incrementally. The text conditioner in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py) demonstrates this pattern, preserving token embeddings across steps as a `StatefulModule`.

## Practical Implementation

The public API abstracts these internals while exposing the streaming functionality through **`TTSModel`**.

### Basic Streaming Generation

```python
from pocket_tts import TTSModel

# Initialise the model (downloads weights on first use)

tts = TTSModel()

# Generate audio from a text prompt – the method yields successive audio chunks

for audio_chunk in tts.generate_audio_stream(
    text="Hello world, this is streaming text-to-speech."
):
    # audio_chunk is a NumPy array (float32) containing the next 80 ms of audio

    play(audio_chunk)  # Stream to audio device or network

```

### Manual State Reset

Interrupt ongoing generation or prepare for a new utterance by clearing all KV caches:

```python
tts.reset_state()   # Clears all KV-caches inside StatefulModules

```

### Voice Cloning with Cached States

When performing voice cloning, the cached voice state is managed separately via an LRU cache, but the underlying KV-cache logic remains identical:

```python
voice_state = tts.get_state_for_audio_prompt("my_voice.wav")
tts.generate_audio_stream(text="Cloned voice speaking.", voice_state=voice_state)

```

## Key Source Files

The streaming architecture is implemented across these core files:

- **[`pocket_tts/modules/stateful_module.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/stateful_module.py)**: Defines the `StatefulModule` base class with `reset_state()` and the `self._state` dictionary for KV-cache storage.
- **[`pocket_tts/modules/transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/transformer.py)**: Implements `StreamingTransformer` and `StreamingMultiheadAttention` that perform incremental attention using the KV cache.
- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Orchestrates the streaming pipeline, calling `reset_state()` at utterance boundaries and driving `generate_audio_stream()`.
- **[`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py)**: Demonstrates stateful conditioning, preserving token embeddings across generation steps.

## Summary

- Pocket-TTS generates audio **frame-by-frame** (every 80 ms) to enable real-time streaming.
- **`StatefulModule`** provides the foundation for state preservation via `self._state`, allowing tensors to persist across forward calls.
- **KV cache management** stores attention keys and values incrementally, so only new queries are computed at each step rather than full attention matrices.
- The **`reset_state()`** method clears caches between utterances, preventing state leakage.
- This architecture enables **low-latency, CPU-only speech synthesis** by sharing cached context across the text encoder, flow LM, and audio decoder.

## Frequently Asked Questions

### What is the purpose of StatefulModule in Pocket-TTS?

**StatefulModule** serves as the base class for all streaming components in Pocket-TTS, providing a standardized way to preserve internal tensors across successive forward calls. It stores the KV cache in `self._state` and exposes `reset_state()` to clear this memory when starting new utterances.

### How does the KV cache reduce computational overhead?

The KV cache eliminates redundant computation by storing previously calculated attention keys and values. When processing a new frame, the model only computes the query for that specific frame and reuses the cached keys and values from all previous positions, reducing the per-step cost from quadratic to linear relative to sequence length.

### When should I call reset_state() manually?

Call `tts.reset_state()` when interrupting generation mid-utterance, switching between unrelated text prompts, or when you need to guarantee that no residual attention context from previous inputs affects new synthesis. The `generate_audio_stream()` method automatically resets state at utterance boundaries.

### Can the streaming architecture run on CPU-only devices?

Yes, the streaming architecture is specifically designed for **CPU-only operation**. By processing audio frame-by-frame and using the KV cache to avoid recomputing full attention matrices, Pocket-TTS maintains real-time performance without requiring GPU acceleration.