pocket-tts
A TTS that fits in your CPU (and pocket)
Explore Pocket TTS alternative implementations including ONNX Web, WebAssembly, MLX, and C++ ports. Deploy TTS in browsers, on Apple Silicon, and embedded systems.
How to Debug KV Cache Expansion Issues Using `_expand_kv_cache()` in Pocket-TTSDebug KV cache expansion issues in Pocket-TTS by inspecting shapes and NaN values before and after calling _expand_kv_cache(). Trace tensor operations with LoggingMode.
Pocket-tts Performance Characteristics on CPUs: RTF, Latency, and Memory UsageDiscover Pocket-tts performance on CPUs. Explore RTF, latency, and memory usage benchmarks. Achieve 2-6x faster audio generation with low RAM consumption.
How Pocket-TTS Handles Text Tokenization with SentencePiece and LUTConditionerLearn how Pocket-TTS tokenizes text using SentencePiece and LUTConditioner. Discover how raw text becomes dense embedding sequences for efficient model processing.
StreamingTransformer with RoPE: How Pocket-TTS Enables Real-Time Speech SynthesisDiscover how Pocket-TTS uses StreamingTransformer with RoPE for real-time speech synthesis. This custom architecture enables low-latency, incremental audio frame generation for efficient inference.
How the LRU Cache Optimizes Voice Prompts in Pocket-TTS with `_cached_get_state_for_audio_prompt()`Discover how the LRU cache optimizes voice prompts in Pocket-TTS with _cached_get_state_for_audio_prompt. Learn how it avoids redundant downloads and computations by caching the last two voice prompt states.
How to Use Predefined Voices in Pocket‑TTS: A Complete GuideExplore predefined voices in Pocket-TTS with our guide. Learn how to select voices like cosette marius and lola programmatically using TTSModel.get_state_for_audio_prompt.
How to Integrate pocket-tts with FastAPI for HTTP-Based Text-to-Speech GenerationEasily integrate pocket-tts with FastAPI for HTTP-based text-to-speech generation. Utilize the built-in POST /tts endpoint to stream WAV audio directly from text input. Learn more.
Pocket-TTS Sample Rate and Audio Resampling: A Technical GuideLearn how Pocket-TTS generates audio at 24 kHz and uses lightweight convolutional blocks for efficient audio resampling. Understand the Pocket-TTS sample rate and resampling process.
How Temperature and Noise_Clamp Parameters Affect Pocket-TTS Generation QualityDiscover how temperature and noise_clamp parameters impact Pocket-TTS generation quality. Understand the trade-offs between diverse speech and stable output for optimal results.
FlowLMModel Architecture: How Pocket‑TTS Generates Latent Audio RepresentationsDiscover the FlowLMModel architecture learn how Pocket-TTS generates latent audio codes in real-time with Lagrangian Self-Distillation for streaming text-to-speech
How EOS Detection Works in Pocket-TTS: Controlling Speech Termination with eos_threshold and frames_after_eosUnderstand EOS detection in Pocket-TTS. Learn how eos_threshold and frames_after_eos control natural speech termination for accurate audio generation. Optimize your TTS.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →