How to Integrate pocket-tts with FastAPI for HTTP-Based Text-to-Speech Generation

The pocket-tts repository includes a built-in FastAPI server that exposes a POST /tts endpoint, streaming WAV audio directly from text input while managing model lifecycle and voice conditioning automatically.

The kyutai-labs/pocket-tts repository provides a production-ready HTTP interface for neural text-to-speech generation. By leveraging the core TTSModel class and FastAPI's streaming response capabilities, you can deploy a low-latency TTS service without writing temporary files or managing complex model state manually.

Architecture Overview

Model Lifecycle Management

In pocket_tts/main.py, the server initializes a global tts_model instance lazily when the serve CLI command executes. According to the source code, the model loads via TTSModel.load_model(...) and persists in memory for the entire process lifetime, eliminating cold-start latency on subsequent requests.

FastAPI Application Structure

The FastAPI object is instantiated at line 46 of pocket_tts/main.py and configured with CORS middleware to support cross-origin browser requests. This architecture separates HTTP handling (in the main module) from generation logic (in pocket_tts/models/tts_model.py), making the codebase maintainable and testable.

The /tts Endpoint Implementation

The POST /tts endpoint accepts multipart/form-data requests with three fields:

  • text (required): The input string to synthesize into speech
  • voice_url (optional): A built-in voice identifier (e.g., "alba")
  • voice_wav (optional): An uploaded WAV file for zero-shot voice cloning

The endpoint validates inputs and resolves voice prompts either from built-in URLs or uploaded audio files. It then calls tts_model.generate_audio_stream(...), which yields raw audio chunks consumed by FastAPI's StreamingResponse with MIME type audio/wav. This streaming approach enables playback to begin before generation completes.

Starting the FastAPI Server

You can launch the server using the provided CLI wrapper or invoke Uvicorn directly against the application object.

Using the CLI (recommended):

uv run pocket-tts serve --host 0.0.0.0 --port 8000

Direct Uvicorn execution:

uvicorn pocket_tts.main:web_app --host 0.0.0.0 --port 8000

The CLI entry point in pocket_tts/__main__.py forwards to pocket_tts.main:serve, which handles default parameter loading from pocket_tts/default_parameters.py before starting the HTTP server.

Client Integration Examples

The following examples demonstrate how to consume the streaming WAV output from Python scripts and command-line tools.

Python client with requests:

import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Hello, world! This is pocket‑tts."),
    # "voice_url": (None, "alba"),  # Optional: use built-in voice

    # "voice_wav": ("voice.wav", open("voice.wav", "rb"))  # Optional: voice clone

}
resp = requests.post(url, files=files, stream=True)

# Stream chunks to disk without loading entire file into memory

with open("output.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

cURL command:

curl -X POST http://localhost:8000/tts \
  -F "text=FastAPI integration demo" \
  -F "voice_url=alba" \
  --output demo.wav

Key Implementation Files

Understanding the source structure helps with customization and debugging:

Summary

  • Built-in server: pocket-tts ships with a complete FastAPI application requiring no additional boilerplate code.
  • Streaming architecture: Audio generates and transmits chunk-by-chunk via StreamingResponse, minimizing latency and memory usage.
  • Flexible voice input: Supports both built-in voice identifiers (voice_url) and custom WAV uploads (voice_wav) through the same endpoint.
  • Persistent model: The global tts_model loads once at startup and stays resident, ensuring fast response times for subsequent requests.

Frequently Asked Questions

Does the pocket-tts FastAPI server require GPU acceleration?

While the server runs on CPU, GPU acceleration significantly improves generation speed for the underlying neural model. The TTSModel class in pocket_tts/models/tts_model.py automatically utilizes available CUDA devices when present, though the FastAPI HTTP layer itself operates independently of the compute backend.

Can I use custom voice prompts with the HTTP endpoint?

Yes. The /tts endpoint accepts an optional voice_wav file field in multipart requests. Upload a clean 24kHz mono WAV file, and the system uses it as a speaker reference for zero-shot voice cloning without requiring model retraining or server restarts.

What audio format does the /tts endpoint return?

The endpoint returns a StreamingResponse with MIME type audio/wav. The stream contains raw PCM WAV data generated by TTSModel.generate_audio_stream, suitable for direct browser playback via the <audio> element or saving to disk as a standard WAV file.

How do I deploy pocket-tts FastAPI in production?

Deploy using standard ASGI server configurations such as Gunicorn with Uvicorn workers or Hypercorn. Ensure the model loads at startup to prevent timeouts on first requests, and consider implementing request rate limiting since generation is computationally intensive. The global tts_model instance in pocket_tts/main.py is process-safe for single-worker deployments but requires careful consideration when scaling horizontally across multiple workers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →