# Using Voicebox REST API for Programmatic Voice Generation: A Complete Integration Guide

> Integrate Voicebox REST API for programmatic voice generation. This guide details how to programmatically generate speech using FastAPI, SSE, and WAV audio files for automated voice synthesis.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: how-to-guide
- Published: 2026-04-14

---

**Voicebox exposes a FastAPI-based REST interface on port `17493` that queues text-to-speech requests via `POST /generate`, streams real-time progress through Server-Sent Events at `/generate/{id}/status`, and delivers WAV audio files, enabling fully automated voice synthesis without the desktop UI.**

Voicebox is an open-source text-to-speech application that provides a comprehensive REST API for programmatic voice generation. According to the `jamiepine/voicebox` source code, the FastAPI server handles asynchronous GPU task queuing, real-time status monitoring, and audio streaming through well-defined endpoints. This guide covers the exact API contract, request lifecycle, and integration patterns needed to embed Voicebox into automated workflows.

## Understanding the API Architecture

The Voicebox server is built on **FastAPI** with endpoints defined in [`backend/routes/generations.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/generations.py). All data validation uses **Pydantic** models located in [`backend/models.py`](https://github.com/jamiepine/voicebox/blob/main/backend/models.py), including `GenerationRequest` and `GenerationResponse`. 

Long-running TTS jobs are not executed synchronously. Instead, [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) maintains an in-process async queue that serializes GPU work, ensuring only one generation runs at a time while tracking active tasks. When a job completes, [`backend/services/history.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/history.py) updates the SQLite database (defined in [`backend/database/models.py`](https://github.com/jamiepine/voicebox/blob/main/backend/database/models.py)) and pushes status changes to connected SSE clients.

## Submitting Your First Generation Request

To initiate speech synthesis, send a `POST` request to `/generate` with a JSON payload matching the `GenerationRequest` schema. Required fields include `profile_id` (UUID of the voice profile), `text` (the script), and optionally `language`, `engine` (defaults to the profile’s setting or "qwen"), and `normalize` (boolean for audio normalization).

The endpoint immediately returns a `GenerationResponse` containing the generation `id`, initial `status` set to `"generating"`, and the target `audio_path`. The actual inference runs in the background via `task_manager.start_generation` and `enqueue_generation` as implemented in the task queue service.

```bash
curl -X POST http://localhost:17493/generate \
  -H "Content-Type: application/json" \
  -d '{
        "profile_id": "my-profile-uuid",
        "text": "Hello, this is a programmatic voice sample.",
        "language": "en",
        "engine": "qwen",
        "normalize": true
      }' | jq .

```

**Response:**

```json
{
  "id": "9f7c2a12-...",
  "profile_id": "my-profile-uuid",
  "text": "Hello, this is a programmatic voice sample.",
  "language": "en",
  "audio_path": "file:///voicebox/generated/9f7c2a12.wav",
  "duration": 2.34,
  "engine": "qwen",
  "status": "generating",
  "created_at": "2026-04-14T10:12:03.456Z"
}

```

## Monitoring Generation Progress in Real Time

Voicebox offers two methods to track job status: **Server-Sent Events (SSE)** for push updates or simple **HTTP polling**.

### Using Server-Sent Events

Connect to `GET /generate/{id}/status` using an `EventSource`. The endpoint defined in [`backend/routes/generations.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/generations.py) yields JSON payloads whenever [`backend/services/history.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/history.py) updates the generation record.

```javascript
const genId = '9f7c2a12-...';
const evtSource = new EventSource(`http://localhost:17493/generate/${genId}/status`);

evtSource.onmessage = (e) => {
  const payload = JSON.parse(e.data);
  console.log('Status update:', payload.status);
  if (payload.status === 'completed') {
    evtSource.close();
    // Fetch the audio file via payload.audio_path
  }
};

```

### Polling with Python

If SSE is unavailable in your environment, poll the same endpoint with standard `GET` requests. The status progresses through `"loading_model"` → `"generating"` → `"completed"` or `"failed"`.

```python
import time, requests

base = "http://localhost:17493"
gen_id = "9f7c2a12-..."

while True:
    r = requests.get(f"{base}/generate/{gen_id}/status")
    data = r.json()
    print(f"Status: {data['status']}, duration: {data.get('duration')}")
    if data["status"] in ("completed", "failed"):
        break
    time.sleep(1)

```

## Streaming and Downloading Audio Files

Once the status reaches `"completed"`, the `audio_path` field contains a `file://` URI to the WAV file. For applications that cannot wait for disk persistence, Voicebox provides a blocking `POST /generate/stream` endpoint that returns the raw audio bytes immediately after synthesis finishes.

```javascript
const axios = require('axios');
const fs = require('fs');

async function streamWav() {
  const resp = await axios.post(
    'http://localhost:17493/generate/stream',
    {
      profile_id: 'my-profile-uuid',
      text: 'Streaming example.',
      language: 'en'
    },
    {
      responseType: 'stream',
      headers: { 'Content-Type': 'application/json' }
    }
  );

  const out = fs.createWriteStream('output.wav');
  resp.data.pipe(out);
  out.on('finish', () => console.log('WAV saved as output.wav'));
}

streamWav();

```

## Retrying and Regenerating Audio

Voicebox supports non-destructive iteration through **retry** and **regenerate** endpoints.

**Retry** (`POST /generate/{id}/retry`): Re-runs the exact same text and seed, useful for transient GPU failures. This skips the effects pipeline to save processing time.

```bash
curl -X POST http://localhost:17493/generate/9f7c2a12-.../retry

```

**Regenerate** (`POST /generate/{id}/regenerate`): Creates a new variation with a random seed, stored as a separate version file (`generationId_<random>.wav`) labeled `take-N`. This allows A/B testing of voice renderings without losing previous outputs.

```bash
curl -X POST http://localhost:17493/generate/9f7c2a12-.../regenerate

```

Version tracking is managed through the SQLAlchemy models in [`backend/database/models.py`](https://github.com/jamiepine/voicebox/blob/main/backend/database/models.py).

## The Generation Pipeline Deep Dive

Understanding the internal flow helps debug failures and optimize request patterns:

1. **Endpoint validation**: [`backend/routes/generations.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/generations.py) validates the `GenerationRequest` against Pydantic models in [`backend/models.py`](https://github.com/jamiepine/voicebox/blob/main/backend/models.py), fetches the profile, and creates a database record via `history.create_generation`.

2. **Task queuing**: `task_manager.start_generation` in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) registers the job ID and hands the coroutine to `enqueue_generation`, which serializes execution to prevent GPU contention.

3. **Core synthesis**: `run_generation` in [`backend/services/generation.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/generation.py) orchestrates the pipeline:
   - Loads the TTS engine via `load_engine_model`
   - Builds a voice prompt using `profiles.create_voice_prompt_for_profile`
   - Processes text through `generate_chunked` in [`backend/utils/chunked_tts.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/chunked_tts.py), which splits long scripts, applies cross-fading, and optionally trims silence
   - Applies audio effects if specified using the Pedalboard-based pipeline in [`backend/utils/effects.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/effects.py)
   - Normalizes audio amplitude when `normalize=True`
   - Persists the final WAV and creates version records via `_save_generate`

4. **Status broadcasting**: [`backend/services/history.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/history.py) updates the generation status at each phase, which the SSE endpoint reads to push updates to clients.

## Summary

- Voicebox runs a FastAPI server on port `17493` with endpoints for generating, monitoring, and streaming speech
- Submit text via `POST /generate` with a profile ID and receive a generation ID immediately while processing queues asynchronously
- Track progress via Server-Sent Events at `/generate/{id}/status` or poll the SQLite-backed endpoint for status values: `loading_model`, `generating`, `completed`, or `failed`
- Download completed WAV files from the returned `audio_path` or stream bytes directly via `POST /generate/stream`
- Retry failed jobs with the same parameters or create variation takes using the retry and regenerate endpoints
- All operations are backed by SQLite persistence defined in [`backend/database/models.py`](https://github.com/jamiepine/voicebox/blob/main/backend/database/models.py) and an async task queue in [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) that serializes GPU-intensive work

## Frequently Asked Questions

### How do I authenticate requests to the Voicebox API?

The open-source `jamiepine/voicebox` implementation does not include authentication middleware; it assumes local network security. When exposing the server externally, place it behind a reverse proxy (such as Nginx or Traefik) with API key validation or OAuth2.

### What is the maximum text length supported by the API?

The API accepts arbitrary text lengths in the `text` field, but the background service automatically chunks long scripts using the utility in [`backend/utils/chunked_tts.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/chunked_tts.py), applying cross-fades between segments to create seamless audio without requiring manual splitting by the client.

### Can I apply real-time audio effects through the API?

Yes, the `GenerationRequest` accepts effect parameters that [`backend/utils/effects.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/effects.py) processes using Pedalboard after the raw audio is generated. Alternatively, generate a clean version first, then access the effects processing separately or manage multiple versions through the database-backed version system.

### How does the task queue handle concurrent requests?

The [`backend/services/task_queue.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/task_queue.py) implements an in-process async queue that serializes GPU-intensive generation tasks, ensuring only one runs at a time while maintaining a registry of active tasks. Client requests return immediately with a generation ID, and the actual inference runs in the background, allowing the API to accept new requests while previous ones complete.