How to Use the pocket-tts Python API to Generate Speech
The pocket-tts Python API exposes the TTSModel class to handle model loading, voice conditioning, and speech generation through methods like load_model(), get_state_for_audio_prompt(), and generate_audio().
The kyutai-labs/pocket-tts repository delivers a fully open-source text-to-speech system that operates entirely on CPU without requiring external GPU resources. Using the pocket-tts Python API, developers can synthesize natural speech from text or clone voices from audio samples with just a few lines of code. This guide examines the core methods defined in pocket_tts/models/tts_model.py and demonstrates practical implementations for both batch and streaming generation.
Core Architecture Overview
The API combines two neural components to produce audio:
- FlowLMModel (
pocket_tts/models/flow_lm.py): A transformer-based flow language model that converts tokenized text into latent audio codes. - MimiModel (
pocket_tts/models/mimi.py): A neural audio codec that compresses and decompresses audio to and from latent representations.
The TTSModel class orchestrates these components, exposing high-level methods that handle tokenization, latent generation, and audio decoding internally.
Loading the Model with TTSModel.load_model()
To begin, initialize the model using the static load_model() method. This downloads the configuration and model weights from HuggingFace, constructs the FlowLM and Mimi sub-models, and optionally applies int-8 quantization for reduced memory usage.
According to the source code in pocket_tts/models/tts_model.py (lines 33-44), the method handles all setup automatically:
from pocket_tts import TTSModel
# Load the default model (English, CPU‑only)
model = TTSModel.load_model()
By default, this loads the English language model with pre-configured hyperparameters from pocket_tts/default_parameters.py.
Voice Cloning via get_state_for_audio_prompt()
For voice cloning, the API uses get_state_for_audio_prompt() to encode audio prompts into a model state that carries speaker characteristics. This method accepts HuggingFace URLs, local WAV files, or pre-computed .safetensors files.
As implemented in pocket_tts/models/tts_model.py (lines 89-107), the method returns a state object compatible with the generation methods:
# Obtain a voice state from a public voice sample on HF
voice_state = model.get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Or use a local file
voice_state = model.get_state_for_audio_prompt("./my_voice.wav")
If no voice state is provided, the model uses its built-in default voice.
Generating Speech
The API provides two generation modes depending on your latency and memory requirements.
Synchronous Generation with generate_audio()
The generate_audio() method performs full-audio generation and returns a single torch.Tensor with shape [channels, samples]. Internally, it streams latent generation, decodes each latent with Mimi, and concatenates the results.
Reference: pocket_tts/models/tts_model.py (lines 84-105).
# Generate with default voice (using dummy init state)
audio_tensor = model.generate_audio(
model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1), # dummy state
text_to_generate="Hello, world! This is Pocket‑TTS."
)
# `audio_tensor` has shape [channels, samples]; save to a WAV file
import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())
Real-time Streaming with generate_audio_stream()
For real-time applications, generate_audio_stream() yields audio chunks as they become available, enabling playback before the full sequence is generated. This method returns a generator that produces tensors incrementally.
Reference: pocket_tts/models/tts_model.py (lines 124-146).
# Stream the output chunk‑by‑chunk
for i, chunk in enumerate(
model.generate_audio_stream(
model_state=voice_state,
text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
)
):
print(f"Chunk {i}: {chunk.shape[0]} samples")
# Here you could feed `chunk` directly to an audio playback library
Complete Usage Examples
Basic Generation with Default Voice
This example demonstrates synthesis without voice cloning, using the model's built-in English voice:
from pocket_tts import TTSModel
# Load the default model (English, CPU‑only)
model = TTSModel.load_model()
# No voice prompt → uses the built‑in default voice
audio_tensor = model.generate_audio(
model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1), # dummy state
text_to_generate="Hello, world! This is Pocket‑TTS."
)
# `audio_tensor` has shape [channels, samples]; save to a WAV file
import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())
Voice Cloning from HuggingFace
To generate speech in a cloned voice from a public audio sample:
from pocket_tts import TTSModel
model = TTSModel.load_model()
# Obtain a voice state from a public voice sample on HF
voice_state = model.get_state_for_audio_prompt(
"hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)
# Generate speech in the cloned voice
audio = model.generate_audio(
model_state=voice_state,
text_to_generate="Good morning! How are you today?"
)
import scipy.io.wavfile
scipy.io.wavfile.write("cloned.wav", model.sample_rate, audio.numpy())
Real-time Streaming Implementation
For applications requiring low-latency audio playback:
from pocket_tts import TTSModel
model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt(
"./my_voice.wav" # local WAV file you recorded
)
# Stream the output chunk‑by‑chunk
for i, chunk in enumerate(
model.generate_audio_stream(
model_state=voice_state,
text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
)
):
print(f"Chunk {i}: {chunk.shape[0]} samples")
# Here you could feed `chunk` directly to an audio playback library
Key Source Files
pocket_tts/models/tts_model.py: High-level API containing theTTSModelclass,load_model(), and generation methods.pocket_tts/models/flow_lm.py: FlowLM transformer implementation that generates latent audio codes from text tokens.pocket_tts/models/mimi.py: Mimi neural codec for audio compression and decompression.pocket_tts/default_parameters.py: Default hyperparameters including temperature and decode steps.pocket_tts/conditioners/text.py: SentencePiece tokenizer and embedding lookup tables used by FlowLM.pocket_tts/utils/config.py: Pydantic configuration handling for YAML model configs.
Summary
- The
TTSModelclass serves as the primary entry point for the pocket-tts Python API, combining FlowLM and Mimi models for end-to-end synthesis. - Use
TTSModel.load_model()to automatically download weights and initialize the pipeline, with optional int-8 quantization. - Voice cloning requires
get_state_for_audio_prompt(), which accepts HuggingFace URLs, local WAV files, or.safetensorsfiles to capture speaker characteristics. - Batch processing uses
generate_audio()to return complete audio tensors, whilegenerate_audio_stream()provides real-time chunk generation for streaming applications. - All processing occurs on CPU without GPU dependencies, making the library suitable for edge deployment.
Frequently Asked Questions
Does pocket-tts require a GPU for speech generation?
No. According to the kyutai-labs/pocket-tts source code, the entire pipeline—including the FlowLM transformer and Mimi codec—is optimized for CPU inference. The load_model() method configures the models for CPU execution by default, though GPU acceleration may be available depending on your PyTorch installation.
What audio formats are supported for voice cloning?
The get_state_for_audio_prompt() method supports three input types: HuggingFace repository URLs (using the hf:// protocol), local WAV file paths, and pre-encoded .safetensors files containing pre-computed voice states. The method handles audio loading and state encoding automatically.
How do I choose between generate_audio and generate_audio_stream?
Use generate_audio() when you need the complete audio file for saving or post-processing, as it returns a single concatenated tensor. Use generate_audio_stream() for real-time applications where latency matters, as it yields audio chunks incrementally via a Python generator, allowing playback to begin before generation completes.
Where does the library download model weights from?
The TTSModel.load_model() method automatically downloads configuration files and model weights from HuggingFace repositories. The download occurs on first initialization and is cached locally for subsequent uses. You can specify custom model configurations by passing parameters to load_model() as defined in pocket_tts/utils/config.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →