# pocket-tts | kyutai | Knowledge Base | Instagit

A TTS that fits in your CPU (and pocket)

GitHub Stars: 6.6k

Repository: https://github.com/kyutai-labs/pocket-tts

---

## Articles

### [Pocket TTS Alternative Implementations: ONNX, WebAssembly, MLX, and C++ Ports](/kyutai-labs/pocket-tts/what-alternative-implementations-exist-for-pocket-tts)

Explore Pocket TTS alternative implementations including ONNX Web, WebAssembly, MLX, and C++ ports. Deploy TTS in browsers, on Apple Silicon, and embedded systems.

- Tags: deep-dive
- Published: 2026-07-11

### [How to Debug KV Cache Expansion Issues Using `_expand_kv_cache()` in Pocket-TTS](/kyutai-labs/pocket-tts/how-do-i-debug-issues-with-kv-cache-expansion-using-_expand_kv_cache)

Debug KV cache expansion issues in Pocket-TTS by inspecting shapes and NaN values before and after calling _expand_kv_cache(). Trace tensor operations with LoggingMode.

- Tags: how-to-guide
- Published: 2026-07-11

### [Pocket-tts Performance Characteristics on CPUs: RTF, Latency, and Memory Usage](/kyutai-labs/pocket-tts/what-are-the-performance-characteristics-on-different-cpus)

Discover Pocket-tts performance on CPUs. Explore RTF, latency, and memory usage benchmarks. Achieve 2-6x faster audio generation with low RAM consumption.

- Tags: performance
- Published: 2026-07-11

### [How Pocket-TTS Handles Text Tokenization with SentencePiece and LUTConditioner](/kyutai-labs/pocket-tts/how-does-the-model-handle-text-tokenization-with-sentencepiece-and-lutconditioner)

Learn how Pocket-TTS tokenizes text using SentencePiece and LUTConditioner. Discover how raw text becomes dense embedding sequences for efficient model processing.

- Tags: deep-dive
- Published: 2026-07-11

### [StreamingTransformer with RoPE: How Pocket-TTS Enables Real-Time Speech Synthesis](/kyutai-labs/pocket-tts/what-is-the-transformer-architecture-used-streamingtransformer-with-rope)

Discover how Pocket-TTS uses StreamingTransformer with RoPE for real-time speech synthesis. This custom architecture enables low-latency, incremental audio frame generation for efficient inference.

- Tags: deep-dive
- Published: 2026-07-11

### [How the LRU Cache Optimizes Voice Prompts in Pocket-TTS with `_cached_get_state_for_audio_prompt()`](/kyutai-labs/pocket-tts/how-does-the-lru-cache-work-for-voice-prompts-with-_cached_get_state_for_audio_prompt)

Discover how the LRU cache optimizes voice prompts in Pocket-TTS with _cached_get_state_for_audio_prompt. Learn how it avoids redundant downloads and computations by caching the last two voice prompt states.

- Tags: internals
- Published: 2026-07-11

### [How to Use Predefined Voices in Pocket‑TTS: A Complete Guide](/kyutai-labs/pocket-tts/what-are-the-predefined-voices-and-how-do-i-select-them-programmatically)

Explore predefined voices in Pocket-TTS with our guide. Learn how to select voices like cosette marius and lola programmatically using TTSModel.get_state_for_audio_prompt.

- Tags: how-to-guide
- Published: 2026-07-11

### [How to Integrate pocket-tts with FastAPI for HTTP-Based Text-to-Speech Generation](/kyutai-labs/pocket-tts/how-do-i-integrate-pocket-tts-with-fastapi-for-http-based-tts-generation)

Easily integrate pocket-tts with FastAPI for HTTP-based text-to-speech generation. Utilize the built-in POST /tts endpoint to stream WAV audio directly from text input. Learn more.

- Tags: how-to-guide
- Published: 2026-07-11

### [Pocket-TTS Sample Rate and Audio Resampling: A Technical Guide](/kyutai-labs/pocket-tts/what-sample-rate-does-the-model-output-and-how-does-audio-resampling-work)

Learn how Pocket-TTS generates audio at 24 kHz and uses lightweight convolutional blocks for efficient audio resampling. Understand the Pocket-TTS sample rate and resampling process.

- Tags: technical-guide
- Published: 2026-07-11

### [How Temperature and Noise_Clamp Parameters Affect Pocket-TTS Generation Quality](/kyutai-labs/pocket-tts/how-do-temperature-and-noise_clamp-parameters-affect-generation-quality)

Discover how temperature and noise_clamp parameters impact Pocket-TTS generation quality. Understand the trade-offs between diverse speech and stable output for optimal results.

- Tags: deep-dive
- Published: 2026-07-11

### [FlowLMModel Architecture: How Pocket‑TTS Generates Latent Audio Representations](/kyutai-labs/pocket-tts/what-is-the-flowlmmodel-architecture-and-how-does-it-generate-latent-audio-representations)

Discover the FlowLMModel architecture learn how Pocket-TTS generates latent audio codes in real-time with Lagrangian Self-Distillation for streaming text-to-speech

- Tags: architecture
- Published: 2026-07-11

### [How EOS Detection Works in Pocket-TTS: Controlling Speech Termination with eos_threshold and frames_after_eos](/kyutai-labs/pocket-tts/how-does-eos-detection-work-with-eos_threshold-and-frames_after_eos-parameters)

Understand EOS detection in Pocket-TTS. Learn how eos_threshold and frames_after_eos control natural speech termination for accurate audio generation. Optimize your TTS.

- Tags: internals
- Published: 2026-07-11

### [What Is the Mimi Neural Audio Codec and How It Is Used for Encoding and Decoding](/kyutai-labs/pocket-tts/what-is-the-mimi-neural-audio-codec-and-how-is-it-used-for-encoding-decoding)

Discover the Mimi neural audio codec, a powerful system for compressing and reconstructing high-quality audio. Learn how it enables voice cloning and streaming synthesis in Pocket-TTS.

- Tags: deep-dive
- Published: 2026-07-11

### [How Pocket-TTS Handles Long Text Inputs with `split_into_best_sentences()`](/kyutai-labs/pocket-tts/how-does-the-model-handle-long-text-inputs-with-split_into_best_sentences)

Learn how Pocket-TTS uses split_into_best_sentences() to handle long text inputs by intelligently splitting sentences and preserving natural phrasing without exceeding token limits.

- Tags: how-to-guide
- Published: 2026-07-11

### [How Pocket-TTS Server Mode Handles Concurrent Requests and Thread Safety Limitations](/kyutai-labs/pocket-tts/how-does-the-server-mode-handle-concurrent-requests-and-what-are-its-thread-safety-limitations)

Discover how Pocket-TTS server mode handles concurrent requests and its thread safety limitations. Learn about race conditions with a single global model under load.

- Tags: internals
- Published: 2026-07-11

### [24-Layer Models vs Standard Models in Pocket-TTS: Architecture and Performance Differences](/kyutai-labs/pocket-tts/what-are-the-differences-between-24-layer-models-and-standard-models)

Explore 24-layer vs standard Pocket-TTS models. Discover architectural differences, performance gains, and trade-offs in inference time and memory for superior audio quality.

- Tags: architecture
- Published: 2026-07-11

### [How to Export and Reload Voice States Using Safetensors for Fast Voice Loading in Pocket-TTS](/kyutai-labs/pocket-tts/how-do-i-export-and-reload-voice-states-using-safetensors-for-fast-voice-loading)

Export and reload Voice States with safetensors in Pocket-TTS for instant, fast voice cloning. Bypass encoding with export_model_state and get_state_for_audio_prompt.

- Tags: how-to-guide
- Published: 2026-07-11

### [What is Lagrangian Self Distillation (LSD) and How Does `lsd_decode_steps` Affect Audio Quality](/kyutai-labs/pocket-tts/what-is-lagrangian-self-distillation-lsd-and-how-does-lsd_decode_steps-affect-quality)

Discover Lagrangian Self Distillation (LSD) in pocket-tts. Learn how lsd_decode_steps impacts audio quality for clearer speech or faster generation.

- Tags: deep-dive
- Published: 2026-07-11

### [How Pocket-TTS Streaming Architecture Works with StatefulModule and KV Cache Management](/kyutai-labs/pocket-tts/how-does-the-streaming-architecture-work-with-statefulmodule-and-kv-cache-management)

Understand Pocket-TTS streaming architecture with StatefulModule and KV cache. Achieve real-time, CPU-only speech synthesis by preserving attention keys and values, eliminating redundant computation.

- Tags: architecture
- Published: 2026-07-11

### [Pocket TTS CLI Commands: generate, serve, and export-voice Options Explained](/kyutai-labs/pocket-tts/what-are-the-available-cli-commands-and-their-parameters-generate-serve-export-voice)

Master Pocket TTS CLI commands like generate serve and export_voice to create WAV audio stream audio or export voice states Learn to load models and process voice prompts efficiently

- Tags: cli-tool
- Published: 2026-07-11

### [How to Use apply_dynamic_int8 for Int8 Quantization in Pocket-TTS](/kyutai-labs/pocket-tts/how-can-i-use-int8-quantization-to-reduce-memory-usage-with-apply_dynamic_int8)

Reduce Pocket-TTS memory usage with apply_dynamic_int8. Convert transformer weights to int8 for about 50% memory savings with minimal audio quality loss. Learn how now.

- Tags: how-to-guide
- Published: 2026-07-11

### [How Voice Cloning Works in pocket‑tts Using `get_state_for_audio_prompt()`](/kyutai-labs/pocket-tts/how-does-voice-cloning-work-in-pocket-tts-with-get_state_for_audio_prompt)

Discover how pocket-tts voice cloning uses get_state_for_audio_prompt to extract a compact VoiceState from audio for efficient streaming generation.

- Tags: deep-dive
- Published: 2026-07-11

### [Difference Between generate_audio() and generate_audio_stream() in Pocket-TTS](/kyutai-labs/pocket-tts/what-is-the-difference-between-generate_audio-and-generate_audio_stream-methods)

Understand Pocket-TTS generate_audio() vs generate_audio_stream(). Get full audio tensors or progressive generator chunks for real-time playback with low latency.

- Tags: deep-dive
- Published: 2026-07-11

### [How TTSModel.load_model() Handles Language Configuration and Model Selection in Pocket TTS](/kyutai-labs/pocket-tts/how-does-ttsmodel-load-model-handle-language-configuration-and-what-language-options-are-available)

Discover how TTSModel load model handles language configuration and available options in Pocket TTS. Learn about language identifiers, YAML mappings, and fallback mechanisms.

- Tags: how-to-guide
- Published: 2026-07-11

### [How SentencePiece Tokenization Works in Pocket-TTS LUTConditioner](/kyutai-labs/pocket-tts/sentencepiece-tokenization-pocket-tts-lutconditioner)

Understand SentencePiece tokenization in Pocket-TTS LUTConditioner. Learn how it maps text to integer IDs for embedding vectors with n_bins + 1 lookup table entries.

- Tags: deep-dive
- Published: 2026-07-09

### [How to Debug Pocket-TTS Audio Generation Issues: A Complete Troubleshooting Guide](/kyutai-labs/pocket-tts/debug-pocket-tts-audio-generation-issues)

Solve Pocket-TTS audio generation problems. This guide helps you debug by checking model weights, prompt encoding, and using debug tools. Get clear audio output.

- Tags: how-to-guide
- Published: 2026-07-09

### [How the Multi-Threaded Architecture Accelerates Latent Generation in Pocket-TTS](/kyutai-labs/pocket-tts/pocket-tts-multithreaded-architecture-latent-generation)

Discover how Pocket-TTS multi-threaded architecture accelerates latent generation through parallel processing. Achieve real-time speech synthesis with pipeline parallelism.

- Tags: deep-dive
- Published: 2026-07-09

### [How Pocket-TTS Uses an LRU Cache to Manage Voice Prompt State](/kyutai-labs/pocket-tts/lru-cache-voice-prompt-state-pocket-tts)

Discover how Pocket-TTS optimizes voice prompt state management with an LRU cache. Learn how it slashes redundant audio processing and KV-cache initialization for faster performance.

- Tags: internals
- Published: 2026-07-09

### [How Pocket‑TTS Handles Chunking and Sentence Splitting for Long Texts](/kyutai-labs/pocket-tts/pocket-tts-chunking-sentence-splitting-long-texts)

Discover how Pocket-TTS effectively handles long texts with its smart chunking and sentence splitting strategies. Learn about primary and secondary delimiters for precise segmenting.

- Tags: internals
- Published: 2026-07-09

### [How Pocket-TTS Handles Infinitely Long Text Inputs: Chunking and Streaming Architecture](/kyutai-labs/pocket-tts/pocket-tts-handle-infinite-text-inputs)

Discover how Pocket-TTS manages long text inputs by chunking and streaming audio sequentially. Learn its architecture for efficient, continuous text-to-speech generation.

- Tags: internals
- Published: 2026-07-09

### [How Int8 Quantization Improves Pocket-TTS Inference Speed: A Technical Breakdown](/kyutai-labs/pocket-tts/int8-quantization-pocket-tts-inference-speed)

Learn how Int8 quantization boosts Pocket-TTS inference speed. Discover how 4x memory reduction and faster GEMM kernels enable up to 40% speedup on modern CPUs.

- Tags: performance
- Published: 2026-07-09

### [How to Interact with the Pocket-TTS HTTP /tts Endpoint: A Complete Guide](/kyutai-labs/pocket-tts/pocket-tts-http-tts-endpoint-interaction)

Use this guide to interact with the pocket-tts HTTP /tts endpoint. Learn how to send text and voice to generate WAV audio streams. Perfect for developers.

- Tags: how-to-guide
- Published: 2026-07-09

### [How to Set Up and Run the Pocket-TTS FastAPI Server](/kyutai-labs/pocket-tts/pocket-tts-fastapi-server-setup)

Learn how to set up and run the Pocket-TTS FastAPI server quickly. Access streaming PCM audio and a web UI with simple CLI commands. Get started today.

- Tags: how-to-guide
- Published: 2026-07-09

### [How to Use the pocket-tts CLI for Speech Generation with Custom Options](/kyutai-labs/pocket-tts/pocket-tts-cli-generate-custom-options)

Learn to use the pocket-tts CLI for powerful speech generation. Explore custom voices, parameters, and device settings for high quality text to speech synthesis.

- Tags: how-to-guide
- Published: 2026-07-09

### [How to Load Models Using TTSModel.load_model() in Pocket‑TTS](/kyutai-labs/pocket-tts/ttsmodel-load-model-pocket-tts-usage)

Learn how to load models with TTSModel.load_model() in Pocket-TTS. This method downloads weights and optionally applies int-8 quantization for faster CPU inference.

- Tags: how-to-guide
- Published: 2026-07-09

### [How to Use the pocket-tts Python API to Generate Speech](/kyutai-labs/pocket-tts/pocket-tts-python-api-generate-speech)

Learn how to use the pocket-tts Python API to generate speech. Explore TTSModel methods for loading, conditioning, and creating audio prompts efficiently.

- Tags: how-to-guide
- Published: 2026-07-09

### [How the Mimi Codec Functions in Pocket‑TTS Voice Cloning](/kyutai-labs/pocket-tts/mimi-codec-function-pocket-tts-voice-cloning)

Discover how the Mimi codec functions in Pocket-TTS voice cloning. Learn how it compresses audio into latent representations to match original speaker voices.

- Tags: deep-dive
- Published: 2026-07-09

### [How Pocket-TTS Achieves Voice Cloning with Audio Encoding: A Technical Deep Dive](/kyutai-labs/pocket-tts/pocket-tts-voice-cloning-audio-encoding)

Discover how Pocket-TTS achieves voice cloning using audio encoding via Mimi neural codec and Flow-LM. Learn the technical details for efficient voice generation.

- Tags: deep-dive
- Published: 2026-07-09

### [Understanding StatefulModule in Pocket‑TTS Streaming: The Architecture Behind Real‑Time Voice Synthesis](/kyutai-labs/pocket-tts/statefulmodule-role-pocket-tts-streaming)

Discover the role of StatefulModule in pocket-tts streaming. This abstract base class standardizes temporal state management for real-time voice synthesis, ensuring coherent audio generation.

- Tags: architecture
- Published: 2026-07-09

