How to Set Up and Run the Pocket-TTS FastAPI Server

Pocket-TTS provides a lightweight FastAPI server via the pocket-tts serve CLI command that exposes a /generate endpoint for streaming PCM audio and a web UI at the root path.

The kyutai-labs/pocket-tts repository ships a CPU-optimized HTTP API that wraps the Flow-LM and Mimi codec models, allowing you to synthesize speech through simple HTTP requests. The server implementation is stateful and handles one request at a time, making it ideal for local development and lightweight deployments. This guide covers the complete setup process, CLI usage, and the underlying source code architecture.

Prerequisites and Environment Setup

Before launching the pocket-tts FastAPI server, ensure your environment meets the following requirements:

  • Python 3.10–3.14: The project requires a modern Python version for compatibility with the underlying PyTorch and audio processing dependencies.
  • uv package manager: Install uv via pip install uv or follow the official Astral installation guide. This tool manages dependencies and virtual environments.
  • System dependencies: A working C++ compiler may be required for audio codec compilation during the first setup.

Create the virtual environment and install locked dependencies by running:

uv sync

This command reads pyproject.toml and installs exact versions of PyTorch (CPU), FastAPI, Uvicorn, and the pocket-tts package itself.

Starting the Pocket-TTS Server

Launch the FastAPI application using the built-in CLI entry point defined in pocket_tts/main.py:

uv run pocket-tts serve

By default, the server binds to http://127.0.0.1:8000. To expose the server on all interfaces or change the port, use the --host and --port flags:

uv run pocket-tts serve --host 0.0.0.0 --port 8080

On first startup, the TTSModel.load_model() method automatically downloads the model checkpoint and tokenizer from HuggingFace, caching them locally for subsequent runs.

API Endpoints and Usage

The pocket-tts FastAPI server exposes two primary endpoints:

Root Endpoint (GET /)

Accessing the root path returns a minimal HTML interface served from the pocket_tts/static/ directory. Open http://127.0.0.1:8000/ in a browser to access a web UI where you can type text, adjust generation parameters like temperature and speed, and download synthesized audio without writing client code.

Generation Endpoint (POST /generate)

The /generate endpoint accepts JSON payloads and streams back raw PCM audio as audio/wav. Send a POST request with the following structure:

curl -X POST "http://127.0.0.1:8000/generate" \
     -H "Content-Type: application/json" \
     -d '{"text":"Hello, Pocket TTS!"}' \
     --output hello.wav

The endpoint supports optional parameters including temperature, top_k, and speed to control the generation behavior.

Programmatic Server Control

You can embed the server directly into Python applications without using the CLI. Import the run_server function from the main module:

from pocket_tts.main import run_server

# Starts the FastAPI server on default host/port

run_server()

For batch processing or integration into existing pipelines, bypass the HTTP layer entirely by using the TTSModel class directly:

from pocket_tts import TTSModel
from pocket_tts.utils.audio import write_wav

model = TTSModel()
model.load_model()  # Downloads weights on first call

audio_chunks = model.generate_audio_stream(
    text="FastAPI server in Pocket-TTS!",
    temperature=0.7,
    top_k=50,
)

write_wav("out.wav", audio_chunks)

Architecture and Key Source Files

Understanding the source code structure helps with customization and debugging:

  • pocket_tts/main.py: Contains the CLI entry point (generate, serve subcommands), FastAPI app initialization, and static file mounting logic.
  • pocket_tts/models/tts_model.py: Implements the core TTSModel class responsible for model loading, Flow-LM inference, and the generate_audio_stream() method that yields audio frames.
  • pocket_tts/static/: Directory containing the HTML, JavaScript, and CSS assets for the built-in web interface.
  • pocket_tts/utils/weights_loading.py: Handles hf:// URL resolution and local caching of downloaded model weights.

The server runs exclusively on CPU using a single Torch thread for deterministic performance. Because the implementation maintains model state in memory, concurrent requests are not supported—the server processes one synthesis job at a time.

Summary

  • Installation: Use uv sync to install dependencies after cloning the repository.
  • Launch: Run uv run pocket-tts serve to start the FastAPI server on 127.0.0.1:8000.
  • Endpoints: Access the web UI at / and generate speech via POST requests to /generate.
  • Architecture: The server uses TTSModel.load_model() for initialization and generate_audio_stream() for audio synthesis, with all routes defined in pocket_tts/main.py.
  • Limitations: CPU-only, single-threaded, and stateful—handles one request at a time.

Frequently Asked Questions

Can the pocket-tts FastAPI server handle multiple concurrent requests?

No, the server is stateful and processes one request at a time. The TTSModel instance maintains loaded weights and generation state in memory, so concurrent requests are not supported. For high-throughput scenarios, deploy multiple instances behind a load balancer.

Does the server support GPU acceleration?

No, the pocket-tts FastAPI server runs on CPU by default and is optimized for CPU inference. The model loading and audio generation in TTSModel use PyTorch CPU tensors, ensuring compatibility across hardware without requiring CUDA drivers.

How do I change the default host and port?

Pass the --host and --port arguments to the serve command: uv run pocket-tts serve --host 0.0.0.0 --port 8080. These flags are processed by the CLI entry point in pocket_tts/main.py and passed directly to the Uvicorn server configuration.

Where are the model weights stored after the first download?

Weights are cached locally after the initial TTSModel.load_model() call. The pocket_tts/utils/weights_loading.py module handles hf:// URL resolution and manages the local cache directory, ensuring subsequent server starts do not require re-downloading from HuggingFace.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →