How Pocket-TTS Streaming Architecture Works with StatefulModule and KV Cache Management

Pocket-TTS achieves real-time, CPU-only speech synthesis by generating audio frame-by-frame using StatefulModule to preserve attention keys and values across steps, eliminating redundant computation through incremental KV-cache updates.

Pocket-TTS (kyutai-labs/pocket-tts) implements a frame-level streaming architecture that enables low-latency text-to-speech on consumer hardware. By leveraging StatefulModule and KV cache management, the system avoids recomputing full attention matrices for every audio frame, drastically reducing memory usage and computational overhead during generation.

Frame-by-Frame Generation at 12.5 Hz

Pocket-TTS generates speech frame-by-frame at approximately 12.5 Hz (every 80 ms). Rather than processing entire sequences in a single forward pass, the model operates incrementally, producing small audio chunks that can be streamed to output devices with minimal latency. This approach requires maintaining context across successive forward calls without reloading the full sequence history.

StatefulModule: Centralizing State Preservation

The foundation of the streaming architecture is StatefulModule, a base class defined in pocket_tts/modules/stateful_module.py. This class centralizes two critical responsibilities: preserving internal tensors across forward calls and managing the KV cache for attention mechanisms.

The State Dictionary

Each module inherits a self._state dictionary that persists any internal tensors required for subsequent generation steps. When a transformer block processes a new token or audio frame, it stores computed keys and values in this state container. According to the source code in pocket_tts/modules/stateful_module.py, this dictionary acts as the persistent memory for the module's streaming context.

Resetting State Between Utterances

When generating a new utterance or interrupting current synthesis, the reset_state() method clears all cached tensors. The TTSModel.generate_audio_stream() method automatically triggers this reset at utterance boundaries, ensuring that each new speech sequence starts from a clean slate without residual attention states from previous inputs.

KV Cache Management in Streaming Transformers

The KV cache (key-value cache) mechanism lives entirely within the StatefulModule state system. Transformer-based components like StreamingTransformer and StreamingMultiheadAttention (implemented in pocket_tts/modules/transformer.py) rely on this cache to avoid recomputing attention weights for all previous positions at every step.

Incremental Cache Updates

The cache follows a three-phase update pattern:

  1. First step: The cache is empty. The module computes keys and values from the initial input token or frame and stores them in self._state.

  2. Subsequent steps: The module computes only the new query vector, then concatenates it with cached keys and values from previous steps. The updated cache is stored back into self._state.

  3. Reset: When reset_state() is called (either manually or automatically at utterance boundaries), the cached tensors are cleared.

This incremental approach means the attention layer only computes queries for the fresh token while reusing cached keys and values for all previous positions, drastically reducing per-step computational cost.

Shared Cache Across the Pipeline

Because the KV cache is unified across all streaming modules, the entire pipeline—including the text encoder, flow language model, and audio decoder—can operate incrementally. The text conditioner in pocket_tts/conditioners/text.py demonstrates this pattern, preserving token embeddings across steps as a StatefulModule.

Practical Implementation

The public API abstracts these internals while exposing the streaming functionality through TTSModel.

Basic Streaming Generation

from pocket_tts import TTSModel

# Initialise the model (downloads weights on first use)

tts = TTSModel()

# Generate audio from a text prompt – the method yields successive audio chunks

for audio_chunk in tts.generate_audio_stream(
    text="Hello world, this is streaming text-to-speech."
):
    # audio_chunk is a NumPy array (float32) containing the next 80 ms of audio

    play(audio_chunk)  # Stream to audio device or network

Manual State Reset

Interrupt ongoing generation or prepare for a new utterance by clearing all KV caches:

tts.reset_state()   # Clears all KV-caches inside StatefulModules

Voice Cloning with Cached States

When performing voice cloning, the cached voice state is managed separately via an LRU cache, but the underlying KV-cache logic remains identical:

voice_state = tts.get_state_for_audio_prompt("my_voice.wav")
tts.generate_audio_stream(text="Cloned voice speaking.", voice_state=voice_state)

Key Source Files

The streaming architecture is implemented across these core files:

Summary

  • Pocket-TTS generates audio frame-by-frame (every 80 ms) to enable real-time streaming.
  • StatefulModule provides the foundation for state preservation via self._state, allowing tensors to persist across forward calls.
  • KV cache management stores attention keys and values incrementally, so only new queries are computed at each step rather than full attention matrices.
  • The reset_state() method clears caches between utterances, preventing state leakage.
  • This architecture enables low-latency, CPU-only speech synthesis by sharing cached context across the text encoder, flow LM, and audio decoder.

Frequently Asked Questions

What is the purpose of StatefulModule in Pocket-TTS?

StatefulModule serves as the base class for all streaming components in Pocket-TTS, providing a standardized way to preserve internal tensors across successive forward calls. It stores the KV cache in self._state and exposes reset_state() to clear this memory when starting new utterances.

How does the KV cache reduce computational overhead?

The KV cache eliminates redundant computation by storing previously calculated attention keys and values. When processing a new frame, the model only computes the query for that specific frame and reuses the cached keys and values from all previous positions, reducing the per-step cost from quadratic to linear relative to sequence length.

When should I call reset_state() manually?

Call tts.reset_state() when interrupting generation mid-utterance, switching between unrelated text prompts, or when you need to guarantee that no residual attention context from previous inputs affects new synthesis. The generate_audio_stream() method automatically resets state at utterance boundaries.

Can the streaming architecture run on CPU-only devices?

Yes, the streaming architecture is specifically designed for CPU-only operation. By processing audio frame-by-frame and using the KV cache to avoid recomputing full attention matrices, Pocket-TTS maintains real-time performance without requiring GPU acceleration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →