How to Use the pocket-tts Python API to Generate Speech

The pocket-tts Python API exposes the TTSModel class to handle model loading, voice conditioning, and speech generation through methods like load_model(), get_state_for_audio_prompt(), and generate_audio().

The kyutai-labs/pocket-tts repository delivers a fully open-source text-to-speech system that operates entirely on CPU without requiring external GPU resources. Using the pocket-tts Python API, developers can synthesize natural speech from text or clone voices from audio samples with just a few lines of code. This guide examines the core methods defined in pocket_tts/models/tts_model.py and demonstrates practical implementations for both batch and streaming generation.

Core Architecture Overview

The API combines two neural components to produce audio:

  • FlowLMModel (pocket_tts/models/flow_lm.py): A transformer-based flow language model that converts tokenized text into latent audio codes.
  • MimiModel (pocket_tts/models/mimi.py): A neural audio codec that compresses and decompresses audio to and from latent representations.

The TTSModel class orchestrates these components, exposing high-level methods that handle tokenization, latent generation, and audio decoding internally.

Loading the Model with TTSModel.load_model()

To begin, initialize the model using the static load_model() method. This downloads the configuration and model weights from HuggingFace, constructs the FlowLM and Mimi sub-models, and optionally applies int-8 quantization for reduced memory usage.

According to the source code in pocket_tts/models/tts_model.py (lines 33-44), the method handles all setup automatically:

from pocket_tts import TTSModel

# Load the default model (English, CPU‑only)

model = TTSModel.load_model()

By default, this loads the English language model with pre-configured hyperparameters from pocket_tts/default_parameters.py.

Voice Cloning via get_state_for_audio_prompt()

For voice cloning, the API uses get_state_for_audio_prompt() to encode audio prompts into a model state that carries speaker characteristics. This method accepts HuggingFace URLs, local WAV files, or pre-computed .safetensors files.

As implemented in pocket_tts/models/tts_model.py (lines 89-107), the method returns a state object compatible with the generation methods:


# Obtain a voice state from a public voice sample on HF

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Or use a local file

voice_state = model.get_state_for_audio_prompt("./my_voice.wav")

If no voice state is provided, the model uses its built-in default voice.

Generating Speech

The API provides two generation modes depending on your latency and memory requirements.

Synchronous Generation with generate_audio()

The generate_audio() method performs full-audio generation and returns a single torch.Tensor with shape [channels, samples]. Internally, it streams latent generation, decodes each latent with Mimi, and concatenates the results.

Reference: pocket_tts/models/tts_model.py (lines 84-105).


# Generate with default voice (using dummy init state)

audio_tensor = model.generate_audio(
    model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1),  # dummy state

    text_to_generate="Hello, world! This is Pocket‑TTS."
)

# `audio_tensor` has shape [channels, samples]; save to a WAV file

import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())

Real-time Streaming with generate_audio_stream()

For real-time applications, generate_audio_stream() yields audio chunks as they become available, enabling playback before the full sequence is generated. This method returns a generator that produces tensors incrementally.

Reference: pocket_tts/models/tts_model.py (lines 124-146).


# Stream the output chunk‑by‑chunk

for i, chunk in enumerate(
    model.generate_audio_stream(
        model_state=voice_state,
        text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
    )
):
    print(f"Chunk {i}: {chunk.shape[0]} samples")
    # Here you could feed `chunk` directly to an audio playback library

Complete Usage Examples

Basic Generation with Default Voice

This example demonstrates synthesis without voice cloning, using the model's built-in English voice:

from pocket_tts import TTSModel

# Load the default model (English, CPU‑only)

model = TTSModel.load_model()

# No voice prompt → uses the built‑in default voice

audio_tensor = model.generate_audio(
    model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1),  # dummy state

    text_to_generate="Hello, world! This is Pocket‑TTS."
)

# `audio_tensor` has shape [channels, samples]; save to a WAV file

import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())

Voice Cloning from HuggingFace

To generate speech in a cloned voice from a public audio sample:

from pocket_tts import TTSModel

model = TTSModel.load_model()

# Obtain a voice state from a public voice sample on HF

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Generate speech in the cloned voice

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Good morning! How are you today?"
)

import scipy.io.wavfile
scipy.io.wavfile.write("cloned.wav", model.sample_rate, audio.numpy())

Real-time Streaming Implementation

For applications requiring low-latency audio playback:

from pocket_tts import TTSModel

model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt(
    "./my_voice.wav"          # local WAV file you recorded

)

# Stream the output chunk‑by‑chunk

for i, chunk in enumerate(
    model.generate_audio_stream(
        model_state=voice_state,
        text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
    )
):
    print(f"Chunk {i}: {chunk.shape[0]} samples")
    # Here you could feed `chunk` directly to an audio playback library

Key Source Files

Summary

  • The TTSModel class serves as the primary entry point for the pocket-tts Python API, combining FlowLM and Mimi models for end-to-end synthesis.
  • Use TTSModel.load_model() to automatically download weights and initialize the pipeline, with optional int-8 quantization.
  • Voice cloning requires get_state_for_audio_prompt(), which accepts HuggingFace URLs, local WAV files, or .safetensors files to capture speaker characteristics.
  • Batch processing uses generate_audio() to return complete audio tensors, while generate_audio_stream() provides real-time chunk generation for streaming applications.
  • All processing occurs on CPU without GPU dependencies, making the library suitable for edge deployment.

Frequently Asked Questions

Does pocket-tts require a GPU for speech generation?

No. According to the kyutai-labs/pocket-tts source code, the entire pipeline—including the FlowLM transformer and Mimi codec—is optimized for CPU inference. The load_model() method configures the models for CPU execution by default, though GPU acceleration may be available depending on your PyTorch installation.

What audio formats are supported for voice cloning?

The get_state_for_audio_prompt() method supports three input types: HuggingFace repository URLs (using the hf:// protocol), local WAV file paths, and pre-encoded .safetensors files containing pre-computed voice states. The method handles audio loading and state encoding automatically.

How do I choose between generate_audio and generate_audio_stream?

Use generate_audio() when you need the complete audio file for saving or post-processing, as it returns a single concatenated tensor. Use generate_audio_stream() for real-time applications where latency matters, as it yields audio chunks incrementally via a Python generator, allowing playback to begin before generation completes.

Where does the library download model weights from?

The TTSModel.load_model() method automatically downloads configuration files and model weights from HuggingFace repositories. The download occurs on first initialization and is cached locally for subsequent uses. You can specify custom model configurations by passing parameters to load_model() as defined in pocket_tts/utils/config.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →