# How to Integrate pocket-tts with FastAPI for HTTP-Based Text-to-Speech Generation

> Easily integrate pocket-tts with FastAPI for HTTP-based text-to-speech generation. Utilize the built-in POST /tts endpoint to stream WAV audio directly from text input. Learn more.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-11

---

**The pocket-tts repository includes a built-in FastAPI server that exposes a POST `/tts` endpoint, streaming WAV audio directly from text input while managing model lifecycle and voice conditioning automatically.**

The `kyutai-labs/pocket-tts` repository provides a production-ready HTTP interface for neural text-to-speech generation. By leveraging the core `TTSModel` class and FastAPI's streaming response capabilities, you can deploy a low-latency TTS service without writing temporary files or managing complex model state manually.

## Architecture Overview

### Model Lifecycle Management

In [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py), the server initializes a global `tts_model` instance lazily when the `serve` CLI command executes. According to the source code, the model loads via `TTSModel.load_model(...)` and persists in memory for the entire process lifetime, eliminating cold-start latency on subsequent requests.

### FastAPI Application Structure

The `FastAPI` object is instantiated at line 46 of [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) and configured with CORS middleware to support cross-origin browser requests. This architecture separates HTTP handling (in the main module) from generation logic (in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)), making the codebase maintainable and testable.

## The `/tts` Endpoint Implementation

The POST `/tts` endpoint accepts `multipart/form-data` requests with three fields:

- **`text`** (required): The input string to synthesize into speech
- **`voice_url`** (optional): A built-in voice identifier (e.g., `"alba"`)
- **`voice_wav`** (optional): An uploaded WAV file for zero-shot voice cloning

The endpoint validates inputs and resolves voice prompts either from built-in URLs or uploaded audio files. It then calls `tts_model.generate_audio_stream(...)`, which yields raw audio chunks consumed by FastAPI's `StreamingResponse` with MIME type `audio/wav`. This streaming approach enables playback to begin before generation completes.

## Starting the FastAPI Server

You can launch the server using the provided CLI wrapper or invoke Uvicorn directly against the application object.

**Using the CLI (recommended):**

```bash
uv run pocket-tts serve --host 0.0.0.0 --port 8000

```

**Direct Uvicorn execution:**

```bash
uvicorn pocket_tts.main:web_app --host 0.0.0.0 --port 8000

```

The CLI entry point in [`pocket_tts/__main__.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/__main__.py) forwards to `pocket_tts.main:serve`, which handles default parameter loading from [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py) before starting the HTTP server.

## Client Integration Examples

The following examples demonstrate how to consume the streaming WAV output from Python scripts and command-line tools.

**Python client with requests:**

```python
import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Hello, world! This is pocket‑tts."),
    # "voice_url": (None, "alba"),  # Optional: use built-in voice

    # "voice_wav": ("voice.wav", open("voice.wav", "rb"))  # Optional: voice clone

}
resp = requests.post(url, files=files, stream=True)

# Stream chunks to disk without loading entire file into memory

with open("output.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

```

**cURL command:**

```bash
curl -X POST http://localhost:8000/tts \
  -F "text=FastAPI integration demo" \
  -F "voice_url=alba" \
  --output demo.wav

```

## Key Implementation Files

Understanding the source structure helps with customization and debugging:

- **[`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py)**: FastAPI server definition, CORS middleware, and `/tts` endpoint routing
- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Core `TTSModel` class handling model loading, voice conditioning, and the `generate_audio_stream` method
- **[`pocket_tts/__main__.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/__main__.py)**: Public CLI entry point that forwards to the server
- **[`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py)**: Default generation hyperparameters (temperature, decode steps)
- **[`pocket_tts/data/audio.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio.py)**: Utility functions for streaming raw audio chunks

## Summary

- **Built-in server**: pocket-tts ships with a complete FastAPI application requiring no additional boilerplate code.
- **Streaming architecture**: Audio generates and transmits chunk-by-chunk via `StreamingResponse`, minimizing latency and memory usage.
- **Flexible voice input**: Supports both built-in voice identifiers (`voice_url`) and custom WAV uploads (`voice_wav`) through the same endpoint.
- **Persistent model**: The global `tts_model` loads once at startup and stays resident, ensuring fast response times for subsequent requests.

## Frequently Asked Questions

### Does the pocket-tts FastAPI server require GPU acceleration?

While the server runs on CPU, GPU acceleration significantly improves generation speed for the underlying neural model. The `TTSModel` class in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) automatically utilizes available CUDA devices when present, though the FastAPI HTTP layer itself operates independently of the compute backend.

### Can I use custom voice prompts with the HTTP endpoint?

Yes. The `/tts` endpoint accepts an optional `voice_wav` file field in multipart requests. Upload a clean 24kHz mono WAV file, and the system uses it as a speaker reference for zero-shot voice cloning without requiring model retraining or server restarts.

### What audio format does the `/tts` endpoint return?

The endpoint returns a `StreamingResponse` with MIME type `audio/wav`. The stream contains raw PCM WAV data generated by `TTSModel.generate_audio_stream`, suitable for direct browser playback via the `<audio>` element or saving to disk as a standard WAV file.

### How do I deploy pocket-tts FastAPI in production?

Deploy using standard ASGI server configurations such as Gunicorn with Uvicorn workers or Hypercorn. Ensure the model loads at startup to prevent timeouts on first requests, and consider implementing request rate limiting since generation is computationally intensive. The global `tts_model` instance in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) is process-safe for single-worker deployments but requires careful consideration when scaling horizontally across multiple workers.