# How to Use the pocket-tts Python API to Generate Speech

> Learn how to use the pocket-tts Python API to generate speech. Explore TTSModel methods for loading, conditioning, and creating audio prompts efficiently.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-09

---

**The pocket-tts Python API exposes the `TTSModel` class to handle model loading, voice conditioning, and speech generation through methods like `load_model()`, `get_state_for_audio_prompt()`, and `generate_audio()`.**

The **kyutai-labs/pocket-tts** repository delivers a fully open-source text-to-speech system that operates entirely on CPU without requiring external GPU resources. Using the pocket-tts Python API, developers can synthesize natural speech from text or clone voices from audio samples with just a few lines of code. This guide examines the core methods defined in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) and demonstrates practical implementations for both batch and streaming generation.

## Core Architecture Overview

The API combines two neural components to produce audio:

- **FlowLMModel** ([`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py)): A transformer-based flow language model that converts tokenized text into latent audio codes.
- **MimiModel** ([`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)): A neural audio codec that compresses and decompresses audio to and from latent representations.

The `TTSModel` class orchestrates these components, exposing high-level methods that handle tokenization, latent generation, and audio decoding internally.

## Loading the Model with TTSModel.load_model()

To begin, initialize the model using the static `load_model()` method. This downloads the configuration and model weights from HuggingFace, constructs the FlowLM and Mimi sub-models, and optionally applies int-8 quantization for reduced memory usage.

According to the source code in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 33-44), the method handles all setup automatically:

```python
from pocket_tts import TTSModel

# Load the default model (English, CPU‑only)

model = TTSModel.load_model()

```

By default, this loads the English language model with pre-configured hyperparameters from [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py).

## Voice Cloning via get_state_for_audio_prompt()

For voice cloning, the API uses `get_state_for_audio_prompt()` to encode audio prompts into a model state that carries speaker characteristics. This method accepts HuggingFace URLs, local WAV files, or pre-computed `.safetensors` files.

As implemented in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 89-107), the method returns a state object compatible with the generation methods:

```python

# Obtain a voice state from a public voice sample on HF

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Or use a local file

voice_state = model.get_state_for_audio_prompt("./my_voice.wav")

```

If no voice state is provided, the model uses its built-in default voice.

## Generating Speech

The API provides two generation modes depending on your latency and memory requirements.

### Synchronous Generation with generate_audio()

The `generate_audio()` method performs full-audio generation and returns a single `torch.Tensor` with shape `[channels, samples]`. Internally, it streams latent generation, decodes each latent with Mimi, and concatenates the results.

Reference: [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 84-105).

```python

# Generate with default voice (using dummy init state)

audio_tensor = model.generate_audio(
    model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1),  # dummy state

    text_to_generate="Hello, world! This is Pocket‑TTS."
)

# `audio_tensor` has shape [channels, samples]; save to a WAV file

import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())

```

### Real-time Streaming with generate_audio_stream()

For real-time applications, `generate_audio_stream()` yields audio chunks as they become available, enabling playback before the full sequence is generated. This method returns a generator that produces tensors incrementally.

Reference: [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 124-146).

```python

# Stream the output chunk‑by‑chunk

for i, chunk in enumerate(
    model.generate_audio_stream(
        model_state=voice_state,
        text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
    )
):
    print(f"Chunk {i}: {chunk.shape[0]} samples")
    # Here you could feed `chunk` directly to an audio playback library

```

## Complete Usage Examples

### Basic Generation with Default Voice

This example demonstrates synthesis without voice cloning, using the model's built-in English voice:

```python
from pocket_tts import TTSModel

# Load the default model (English, CPU‑only)

model = TTSModel.load_model()

# No voice prompt → uses the built‑in default voice

audio_tensor = model.generate_audio(
    model_state=model.flow_lm.init_states(batch_size=1, sequence_length=1),  # dummy state

    text_to_generate="Hello, world! This is Pocket‑TTS."
)

# `audio_tensor` has shape [channels, samples]; save to a WAV file

import scipy.io.wavfile
scipy.io.wavfile.write("hello.wav", model.sample_rate, audio_tensor.numpy())

```

### Voice Cloning from HuggingFace

To generate speech in a cloned voice from a public audio sample:

```python
from pocket_tts import TTSModel

model = TTSModel.load_model()

# Obtain a voice state from a public voice sample on HF

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Generate speech in the cloned voice

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Good morning! How are you today?"
)

import scipy.io.wavfile
scipy.io.wavfile.write("cloned.wav", model.sample_rate, audio.numpy())

```

### Real-time Streaming Implementation

For applications requiring low-latency audio playback:

```python
from pocket_tts import TTSModel

model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt(
    "./my_voice.wav"          # local WAV file you recorded

)

# Stream the output chunk‑by‑chunk

for i, chunk in enumerate(
    model.generate_audio_stream(
        model_state=voice_state,
        text_to_generate="This is a long paragraph that will be streamed chunk by chunk."
    )
):
    print(f"Chunk {i}: {chunk.shape[0]} samples")
    # Here you could feed `chunk` directly to an audio playback library

```

## Key Source Files

- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: High-level API containing the `TTSModel` class, `load_model()`, and generation methods.
- **[`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py)**: FlowLM transformer implementation that generates latent audio codes from text tokens.
- **[`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)**: Mimi neural codec for audio compression and decompression.
- **[`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py)**: Default hyperparameters including temperature and decode steps.
- **[`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py)**: SentencePiece tokenizer and embedding lookup tables used by FlowLM.
- **[`pocket_tts/utils/config.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/config.py)**: Pydantic configuration handling for YAML model configs.

## Summary

- **The `TTSModel` class** serves as the primary entry point for the pocket-tts Python API, combining FlowLM and Mimi models for end-to-end synthesis.
- **Use `TTSModel.load_model()`** to automatically download weights and initialize the pipeline, with optional int-8 quantization.
- **Voice cloning** requires `get_state_for_audio_prompt()`, which accepts HuggingFace URLs, local WAV files, or `.safetensors` files to capture speaker characteristics.
- **Batch processing** uses `generate_audio()` to return complete audio tensors, while `generate_audio_stream()` provides real-time chunk generation for streaming applications.
- **All processing occurs on CPU** without GPU dependencies, making the library suitable for edge deployment.

## Frequently Asked Questions

### Does pocket-tts require a GPU for speech generation?

No. According to the kyutai-labs/pocket-tts source code, the entire pipeline—including the FlowLM transformer and Mimi codec—is optimized for CPU inference. The `load_model()` method configures the models for CPU execution by default, though GPU acceleration may be available depending on your PyTorch installation.

### What audio formats are supported for voice cloning?

The `get_state_for_audio_prompt()` method supports three input types: HuggingFace repository URLs (using the `hf://` protocol), local WAV file paths, and pre-encoded `.safetensors` files containing pre-computed voice states. The method handles audio loading and state encoding automatically.

### How do I choose between generate_audio and generate_audio_stream?

Use **`generate_audio()`** when you need the complete audio file for saving or post-processing, as it returns a single concatenated tensor. Use **`generate_audio_stream()`** for real-time applications where latency matters, as it yields audio chunks incrementally via a Python generator, allowing playback to begin before generation completes.

### Where does the library download model weights from?

The `TTSModel.load_model()` method automatically downloads configuration files and model weights from HuggingFace repositories. The download occurs on first initialization and is cached locally for subsequent uses. You can specify custom model configurations by passing parameters to `load_model()` as defined in [`pocket_tts/utils/config.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/config.py).