How to Interact with the Pocket-TTS HTTP /tts Endpoint: A Complete Guide

The pocket-tts HTTP /tts endpoint accepts a multipart POST request containing text and a voice source, then streams a generated WAV audio file back to the client using FastAPI's StreamingResponse.

The pocket-tts HTTP /tts endpoint provides a stateless, streaming text-to-speech service exposed by the kyutai-labs/pocket-tts repository. Implemented in pocket_tts/main.py, this single endpoint handles voice cloning, predefined voices, and real-time audio generation through a simple HTTP interface that accepts either voice URLs or uploaded audio files.

Endpoint Specification and Request Validation

The /tts endpoint is defined in pocket_tts/main.py and enforces strict input validation before processing. According to lines 35-44 of the source code, every request must include a non-empty text parameter and exactly one voice source: either voice_url or voice_wav. If neither voice parameter is provided, the system automatically selects a default voice appropriate for the model's configured language.

The endpoint accepts multipart/form-data requests, making it compatible with standard HTTP clients, web browsers, and command-line tools like cURL.

Voice Selection and Resolution

The pocket-tts HTTP /tts endpoint supports two distinct methods for specifying the target voice, handled by separate code paths in main.py.

Using the voice_url Parameter

The voice_url field accepts three distinct input types, as implemented in lines 45-61:

  • Predefined voice names such as "alba" that map to built-in voices listed in _ORIGINS_OF_PREDEFINED_VOICES (defined in pocket_tts/utils/utils.py)
  • Standard HTTP/HTTPS URLs pointing to remote WAV files
  • HuggingFace URLs using the hf:// scheme for direct repository access

When a voice_url is provided, the endpoint validates the scheme against supported origins, then retrieves the model state via tts_model._cached_get_state_for_audio_prompt, which caches voice embeddings to avoid redundant processing.

Using the voice_wav Parameter

For custom voice cloning, the voice_wav parameter accepts an uploaded audio file. As shown in lines 66-71, the handler saves the uploaded file to a temporary location on disk, then processes it through tts_model.get_state_for_audio_prompt to extract the voice characteristics needed for generation.

Streaming Audio Generation and Response

Once voice resolution completes, the endpoint initiates streaming generation. The implementation in lines 71-82 of main.py passes the resolved model_state and input text to tts_model.generate_audio_stream, which yields audio frames incrementally.

A helper function generate_data_with_state executes the generation in a background thread, feeding audio chunks into a Queue that the main thread consumes. The endpoint returns a FastAPI StreamingResponse with the following characteristics:

  • Content-Type: audio/wav
  • Content-Disposition: attachment; filename=generated_speech.wav
  • Transfer-Encoding: chunked to enable playback while generation continues

This architecture allows clients to begin playing audio before the full file is generated, reducing latency for long text inputs.

Practical Code Examples

cURL: Text with Default Voice

curl -X POST http://localhost:8000/tts \
     -F "text=Hello, Pocket-TTS!" \
     -o generated.wav

cURL: Specifying a Built-in Voice

curl -X POST http://localhost:8000/tts \
     -F "text=Bonjour le monde!" \
     -F "voice_url=alba" \
     -o generated.wav

cURL: Using a Remote Voice URL

curl -X POST http://localhost:8000/tts \
     -F "text=Guten Tag!" \
     -F "voice_url=https://huggingface.co/kyutai/tts-voices/resolve/main/german/standard.wav" \
     -o generated.wav

cURL: Uploading a Custom WAV File

curl -X POST http://localhost:8000/tts \
     -F "text=Hola, ¿cómo estás?" \
     -F "voice_wav=@my_voice.wav;type=audio/wav" \
     -o generated.wav

Python: Using a Predefined Voice

import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Hello from Python!"),
    "voice_url": (None, "alba")
}
resp = requests.post(url, files=files, stream=True)

with open("generated.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

Python: Uploading a Local Voice File

import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Testing custom voice"),
    "voice_wav": ("my_voice.wav", open("my_voice.wav", "rb"), "audio/wav")
}
resp = requests.post(url, files=files, stream=True)

with open("generated.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

Summary

  • The pocket-tts HTTP /tts endpoint in pocket_tts/main.py provides a single POST interface for text-to-speech generation
  • Requests require text and exactly one voice source: voice_url (string) or voice_wav (file upload)
  • Voice URLs support predefined names, HTTP/HTTPS links, and HuggingFace hf:// URLs validated against _ORIGINS_OF_PREDEFINED_VOICES
  • Custom voice files are processed via tts_model.get_state_for_audio_prompt and cached for subsequent requests
  • The endpoint returns a streaming WAV response using Transfer-Encoding: chunked for real-time playback

Frequently Asked Questions

What audio format does the pocket-tts /tts endpoint return?

The endpoint returns raw PCM audio wrapped in a WAV container with Content-Type: audio/wav. The response uses chunked transfer encoding, allowing clients to stream the audio rather than waiting for the complete file to generate.

Can I use a custom voice file with the HTTP endpoint?

Yes. Upload a WAV file using the voice_wav parameter in a multipart/form-data request. The endpoint saves the file temporarily and processes it through tts_model.get_state_for_audio_prompt to extract voice characteristics for cloning.

How does voice caching work in pocket-tts?

When using voice_url with predefined voices or remote URLs, the endpoint calls tts_model._cached_get_state_for_audio_prompt, which caches the model state in memory. This eliminates redundant processing if the same voice is requested multiple times.

What happens if I don't specify a voice in the request?

If neither voice_url nor voice_wav is provided, the endpoint automatically selects a default voice appropriate for the model's configured language, as implemented in the validation logic at lines 35-44 of pocket_tts/main.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →