How to Integrate pocket-tts with FastAPI for HTTP-Based Text-to-Speech Generation
The pocket-tts repository includes a built-in FastAPI server that exposes a POST /tts endpoint, streaming WAV audio directly from text input while managing model lifecycle and voice conditioning automatically.
The kyutai-labs/pocket-tts repository provides a production-ready HTTP interface for neural text-to-speech generation. By leveraging the core TTSModel class and FastAPI's streaming response capabilities, you can deploy a low-latency TTS service without writing temporary files or managing complex model state manually.
Architecture Overview
Model Lifecycle Management
In pocket_tts/main.py, the server initializes a global tts_model instance lazily when the serve CLI command executes. According to the source code, the model loads via TTSModel.load_model(...) and persists in memory for the entire process lifetime, eliminating cold-start latency on subsequent requests.
FastAPI Application Structure
The FastAPI object is instantiated at line 46 of pocket_tts/main.py and configured with CORS middleware to support cross-origin browser requests. This architecture separates HTTP handling (in the main module) from generation logic (in pocket_tts/models/tts_model.py), making the codebase maintainable and testable.
The /tts Endpoint Implementation
The POST /tts endpoint accepts multipart/form-data requests with three fields:
text(required): The input string to synthesize into speechvoice_url(optional): A built-in voice identifier (e.g.,"alba")voice_wav(optional): An uploaded WAV file for zero-shot voice cloning
The endpoint validates inputs and resolves voice prompts either from built-in URLs or uploaded audio files. It then calls tts_model.generate_audio_stream(...), which yields raw audio chunks consumed by FastAPI's StreamingResponse with MIME type audio/wav. This streaming approach enables playback to begin before generation completes.
Starting the FastAPI Server
You can launch the server using the provided CLI wrapper or invoke Uvicorn directly against the application object.
Using the CLI (recommended):
uv run pocket-tts serve --host 0.0.0.0 --port 8000
Direct Uvicorn execution:
uvicorn pocket_tts.main:web_app --host 0.0.0.0 --port 8000
The CLI entry point in pocket_tts/__main__.py forwards to pocket_tts.main:serve, which handles default parameter loading from pocket_tts/default_parameters.py before starting the HTTP server.
Client Integration Examples
The following examples demonstrate how to consume the streaming WAV output from Python scripts and command-line tools.
Python client with requests:
import requests
url = "http://localhost:8000/tts"
files = {
"text": (None, "Hello, world! This is pocket‑tts."),
# "voice_url": (None, "alba"), # Optional: use built-in voice
# "voice_wav": ("voice.wav", open("voice.wav", "rb")) # Optional: voice clone
}
resp = requests.post(url, files=files, stream=True)
# Stream chunks to disk without loading entire file into memory
with open("output.wav", "wb") as f:
for chunk in resp.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
cURL command:
curl -X POST http://localhost:8000/tts \
-F "text=FastAPI integration demo" \
-F "voice_url=alba" \
--output demo.wav
Key Implementation Files
Understanding the source structure helps with customization and debugging:
pocket_tts/main.py: FastAPI server definition, CORS middleware, and/ttsendpoint routingpocket_tts/models/tts_model.py: CoreTTSModelclass handling model loading, voice conditioning, and thegenerate_audio_streammethodpocket_tts/__main__.py: Public CLI entry point that forwards to the serverpocket_tts/default_parameters.py: Default generation hyperparameters (temperature, decode steps)pocket_tts/data/audio.py: Utility functions for streaming raw audio chunks
Summary
- Built-in server: pocket-tts ships with a complete FastAPI application requiring no additional boilerplate code.
- Streaming architecture: Audio generates and transmits chunk-by-chunk via
StreamingResponse, minimizing latency and memory usage. - Flexible voice input: Supports both built-in voice identifiers (
voice_url) and custom WAV uploads (voice_wav) through the same endpoint. - Persistent model: The global
tts_modelloads once at startup and stays resident, ensuring fast response times for subsequent requests.
Frequently Asked Questions
Does the pocket-tts FastAPI server require GPU acceleration?
While the server runs on CPU, GPU acceleration significantly improves generation speed for the underlying neural model. The TTSModel class in pocket_tts/models/tts_model.py automatically utilizes available CUDA devices when present, though the FastAPI HTTP layer itself operates independently of the compute backend.
Can I use custom voice prompts with the HTTP endpoint?
Yes. The /tts endpoint accepts an optional voice_wav file field in multipart requests. Upload a clean 24kHz mono WAV file, and the system uses it as a speaker reference for zero-shot voice cloning without requiring model retraining or server restarts.
What audio format does the /tts endpoint return?
The endpoint returns a StreamingResponse with MIME type audio/wav. The stream contains raw PCM WAV data generated by TTSModel.generate_audio_stream, suitable for direct browser playback via the <audio> element or saving to disk as a standard WAV file.
How do I deploy pocket-tts FastAPI in production?
Deploy using standard ASGI server configurations such as Gunicorn with Uvicorn workers or Hypercorn. Ensure the model loads at startup to prevent timeouts on first requests, and consider implementing request rate limiting since generation is computationally intensive. The global tts_model instance in pocket_tts/main.py is process-safe for single-worker deployments but requires careful consideration when scaling horizontally across multiple workers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →