# How to Interact with the Pocket-TTS HTTP /tts Endpoint: A Complete Guide

> Use this guide to interact with the pocket-tts HTTP /tts endpoint. Learn how to send text and voice to generate WAV audio streams. Perfect for developers.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-09

---

**The pocket-tts HTTP `/tts` endpoint accepts a multipart POST request containing text and a voice source, then streams a generated WAV audio file back to the client using FastAPI's StreamingResponse.**

The pocket-tts HTTP /tts endpoint provides a stateless, streaming text-to-speech service exposed by the `kyutai-labs/pocket-tts` repository. Implemented in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py), this single endpoint handles voice cloning, predefined voices, and real-time audio generation through a simple HTTP interface that accepts either voice URLs or uploaded audio files.

## Endpoint Specification and Request Validation

The `/tts` endpoint is defined in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) and enforces strict input validation before processing. According to lines 35-44 of the source code, every request must include a non-empty `text` parameter and exactly one voice source: either `voice_url` or `voice_wav`. If neither voice parameter is provided, the system automatically selects a default voice appropriate for the model's configured language.

The endpoint accepts `multipart/form-data` requests, making it compatible with standard HTTP clients, web browsers, and command-line tools like cURL.

## Voice Selection and Resolution

The pocket-tts HTTP /tts endpoint supports two distinct methods for specifying the target voice, handled by separate code paths in [`main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/main.py).

### Using the voice_url Parameter

The `voice_url` field accepts three distinct input types, as implemented in lines 45-61:

- **Predefined voice names** such as `"alba"` that map to built-in voices listed in `_ORIGINS_OF_PREDEFINED_VOICES` (defined in [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py))
- **Standard HTTP/HTTPS URLs** pointing to remote WAV files
- **HuggingFace URLs** using the `hf://` scheme for direct repository access

When a `voice_url` is provided, the endpoint validates the scheme against supported origins, then retrieves the model state via `tts_model._cached_get_state_for_audio_prompt`, which caches voice embeddings to avoid redundant processing.

### Using the voice_wav Parameter

For custom voice cloning, the `voice_wav` parameter accepts an uploaded audio file. As shown in lines 66-71, the handler saves the uploaded file to a temporary location on disk, then processes it through `tts_model.get_state_for_audio_prompt` to extract the voice characteristics needed for generation.

## Streaming Audio Generation and Response

Once voice resolution completes, the endpoint initiates streaming generation. The implementation in lines 71-82 of [`main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/main.py) passes the resolved `model_state` and input `text` to `tts_model.generate_audio_stream`, which yields audio frames incrementally.

A helper function `generate_data_with_state` executes the generation in a background thread, feeding audio chunks into a `Queue` that the main thread consumes. The endpoint returns a FastAPI `StreamingResponse` with the following characteristics:

- **Content-Type**: `audio/wav`
- **Content-Disposition**: `attachment; filename=generated_speech.wav`
- **Transfer-Encoding**: `chunked` to enable playback while generation continues

This architecture allows clients to begin playing audio before the full file is generated, reducing latency for long text inputs.

## Practical Code Examples

### cURL: Text with Default Voice

```bash
curl -X POST http://localhost:8000/tts \
     -F "text=Hello, Pocket-TTS!" \
     -o generated.wav

```

### cURL: Specifying a Built-in Voice

```bash
curl -X POST http://localhost:8000/tts \
     -F "text=Bonjour le monde!" \
     -F "voice_url=alba" \
     -o generated.wav

```

### cURL: Using a Remote Voice URL

```bash
curl -X POST http://localhost:8000/tts \
     -F "text=Guten Tag!" \
     -F "voice_url=https://huggingface.co/kyutai/tts-voices/resolve/main/german/standard.wav" \
     -o generated.wav

```

### cURL: Uploading a Custom WAV File

```bash
curl -X POST http://localhost:8000/tts \
     -F "text=Hola, ¿cómo estás?" \
     -F "voice_wav=@my_voice.wav;type=audio/wav" \
     -o generated.wav

```

### Python: Using a Predefined Voice

```python
import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Hello from Python!"),
    "voice_url": (None, "alba")
}
resp = requests.post(url, files=files, stream=True)

with open("generated.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

```

### Python: Uploading a Local Voice File

```python
import requests

url = "http://localhost:8000/tts"
files = {
    "text": (None, "Testing custom voice"),
    "voice_wav": ("my_voice.wav", open("my_voice.wav", "rb"), "audio/wav")
}
resp = requests.post(url, files=files, stream=True)

with open("generated.wav", "wb") as f:
    for chunk in resp.iter_content(chunk_size=8192):
        if chunk:
            f.write(chunk)

```

## Summary

- The pocket-tts HTTP /tts endpoint in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) provides a single POST interface for text-to-speech generation
- Requests require `text` and exactly one voice source: `voice_url` (string) or `voice_wav` (file upload)
- Voice URLs support predefined names, HTTP/HTTPS links, and HuggingFace `hf://` URLs validated against `_ORIGINS_OF_PREDEFINED_VOICES`
- Custom voice files are processed via `tts_model.get_state_for_audio_prompt` and cached for subsequent requests
- The endpoint returns a streaming WAV response using `Transfer-Encoding: chunked` for real-time playback

## Frequently Asked Questions

### What audio format does the pocket-tts /tts endpoint return?

The endpoint returns raw PCM audio wrapped in a WAV container with `Content-Type: audio/wav`. The response uses chunked transfer encoding, allowing clients to stream the audio rather than waiting for the complete file to generate.

### Can I use a custom voice file with the HTTP endpoint?

Yes. Upload a WAV file using the `voice_wav` parameter in a multipart/form-data request. The endpoint saves the file temporarily and processes it through `tts_model.get_state_for_audio_prompt` to extract voice characteristics for cloning.

### How does voice caching work in pocket-tts?

When using `voice_url` with predefined voices or remote URLs, the endpoint calls `tts_model._cached_get_state_for_audio_prompt`, which caches the model state in memory. This eliminates redundant processing if the same voice is requested multiple times.

### What happens if I don't specify a voice in the request?

If neither `voice_url` nor `voice_wav` is provided, the endpoint automatically selects a default voice appropriate for the model's configured language, as implemented in the validation logic at lines 35-44 of [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py).