# How to Set Up and Run the Pocket-TTS FastAPI Server

> Learn how to set up and run the Pocket-TTS FastAPI server quickly. Access streaming PCM audio and a web UI with simple CLI commands. Get started today.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Pocket-TTS provides a lightweight FastAPI server via the `pocket-tts serve` CLI command that exposes a `/generate` endpoint for streaming PCM audio and a web UI at the root path.**

The `kyutai-labs/pocket-tts` repository ships a CPU-optimized HTTP API that wraps the Flow-LM and Mimi codec models, allowing you to synthesize speech through simple HTTP requests. The server implementation is stateful and handles one request at a time, making it ideal for local development and lightweight deployments. This guide covers the complete setup process, CLI usage, and the underlying source code architecture.

## Prerequisites and Environment Setup

Before launching the pocket-tts FastAPI server, ensure your environment meets the following requirements:

- **Python 3.10–3.14**: The project requires a modern Python version for compatibility with the underlying PyTorch and audio processing dependencies.
- **uv package manager**: Install `uv` via `pip install uv` or follow the official Astral installation guide. This tool manages dependencies and virtual environments.
- **System dependencies**: A working C++ compiler may be required for audio codec compilation during the first setup.

Create the virtual environment and install locked dependencies by running:

```bash
uv sync

```

This command reads [`pyproject.toml`](https://github.com/kyutai-labs/pocket-tts/blob/main/pyproject.toml) and installs exact versions of PyTorch (CPU), FastAPI, Uvicorn, and the pocket-tts package itself.

## Starting the Pocket-TTS Server

Launch the FastAPI application using the built-in CLI entry point defined in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py):

```bash
uv run pocket-tts serve

```

By default, the server binds to `http://127.0.0.1:8000`. To expose the server on all interfaces or change the port, use the `--host` and `--port` flags:

```bash
uv run pocket-tts serve --host 0.0.0.0 --port 8080

```

On first startup, the `TTSModel.load_model()` method automatically downloads the model checkpoint and tokenizer from HuggingFace, caching them locally for subsequent runs.

## API Endpoints and Usage

The pocket-tts FastAPI server exposes two primary endpoints:

### Root Endpoint (`GET /`)

Accessing the root path returns a minimal HTML interface served from the `pocket_tts/static/` directory. Open `http://127.0.0.1:8000/` in a browser to access a web UI where you can type text, adjust generation parameters like `temperature` and `speed`, and download synthesized audio without writing client code.

### Generation Endpoint (`POST /generate`)

The `/generate` endpoint accepts JSON payloads and streams back raw PCM audio as `audio/wav`. Send a POST request with the following structure:

```bash
curl -X POST "http://127.0.0.1:8000/generate" \
     -H "Content-Type: application/json" \
     -d '{"text":"Hello, Pocket TTS!"}' \
     --output hello.wav

```

The endpoint supports optional parameters including `temperature`, `top_k`, and `speed` to control the generation behavior.

## Programmatic Server Control

You can embed the server directly into Python applications without using the CLI. Import the `run_server` function from the main module:

```python
from pocket_tts.main import run_server

# Starts the FastAPI server on default host/port

run_server()

```

For batch processing or integration into existing pipelines, bypass the HTTP layer entirely by using the `TTSModel` class directly:

```python
from pocket_tts import TTSModel
from pocket_tts.utils.audio import write_wav

model = TTSModel()
model.load_model()  # Downloads weights on first call

audio_chunks = model.generate_audio_stream(
    text="FastAPI server in Pocket-TTS!",
    temperature=0.7,
    top_k=50,
)

write_wav("out.wav", audio_chunks)

```

## Architecture and Key Source Files

Understanding the source code structure helps with customization and debugging:

- **[`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py)**: Contains the CLI entry point (`generate`, `serve` subcommands), FastAPI app initialization, and static file mounting logic.
- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Implements the core `TTSModel` class responsible for model loading, Flow-LM inference, and the `generate_audio_stream()` method that yields audio frames.
- **`pocket_tts/static/`**: Directory containing the HTML, JavaScript, and CSS assets for the built-in web interface.
- **[`pocket_tts/utils/weights_loading.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/weights_loading.py)**: Handles `hf://` URL resolution and local caching of downloaded model weights.

The server runs exclusively on CPU using a single Torch thread for deterministic performance. Because the implementation maintains model state in memory, concurrent requests are not supported—the server processes one synthesis job at a time.

## Summary

- **Installation**: Use `uv sync` to install dependencies after cloning the repository.
- **Launch**: Run `uv run pocket-tts serve` to start the FastAPI server on `127.0.0.1:8000`.
- **Endpoints**: Access the web UI at `/` and generate speech via POST requests to `/generate`.
- **Architecture**: The server uses `TTSModel.load_model()` for initialization and `generate_audio_stream()` for audio synthesis, with all routes defined in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py).
- **Limitations**: CPU-only, single-threaded, and stateful—handles one request at a time.

## Frequently Asked Questions

### Can the pocket-tts FastAPI server handle multiple concurrent requests?

No, the server is stateful and processes one request at a time. The `TTSModel` instance maintains loaded weights and generation state in memory, so concurrent requests are not supported. For high-throughput scenarios, deploy multiple instances behind a load balancer.

### Does the server support GPU acceleration?

No, the pocket-tts FastAPI server runs on CPU by default and is optimized for CPU inference. The model loading and audio generation in `TTSModel` use PyTorch CPU tensors, ensuring compatibility across hardware without requiring CUDA drivers.

### How do I change the default host and port?

Pass the `--host` and `--port` arguments to the serve command: `uv run pocket-tts serve --host 0.0.0.0 --port 8080`. These flags are processed by the CLI entry point in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) and passed directly to the Uvicorn server configuration.

### Where are the model weights stored after the first download?

Weights are cached locally after the initial `TTSModel.load_model()` call. The [`pocket_tts/utils/weights_loading.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/weights_loading.py) module handles `hf://` URL resolution and manages the local cache directory, ensuring subsequent server starts do not require re-downloading from HuggingFace.