Using Voicebox REST API for Programmatic Voice Generation: A Complete Integration Guide

Voicebox exposes a FastAPI-based REST interface on port 17493 that queues text-to-speech requests via POST /generate, streams real-time progress through Server-Sent Events at /generate/{id}/status, and delivers WAV audio files, enabling fully automated voice synthesis without the desktop UI.

Voicebox is an open-source text-to-speech application that provides a comprehensive REST API for programmatic voice generation. According to the jamiepine/voicebox source code, the FastAPI server handles asynchronous GPU task queuing, real-time status monitoring, and audio streaming through well-defined endpoints. This guide covers the exact API contract, request lifecycle, and integration patterns needed to embed Voicebox into automated workflows.

Understanding the API Architecture

The Voicebox server is built on FastAPI with endpoints defined in backend/routes/generations.py. All data validation uses Pydantic models located in backend/models.py, including GenerationRequest and GenerationResponse.

Long-running TTS jobs are not executed synchronously. Instead, backend/services/task_queue.py maintains an in-process async queue that serializes GPU work, ensuring only one generation runs at a time while tracking active tasks. When a job completes, backend/services/history.py updates the SQLite database (defined in backend/database/models.py) and pushes status changes to connected SSE clients.

Submitting Your First Generation Request

To initiate speech synthesis, send a POST request to /generate with a JSON payload matching the GenerationRequest schema. Required fields include profile_id (UUID of the voice profile), text (the script), and optionally language, engine (defaults to the profile’s setting or "qwen"), and normalize (boolean for audio normalization).

The endpoint immediately returns a GenerationResponse containing the generation id, initial status set to "generating", and the target audio_path. The actual inference runs in the background via task_manager.start_generation and enqueue_generation as implemented in the task queue service.

curl -X POST http://localhost:17493/generate \
  -H "Content-Type: application/json" \
  -d '{
        "profile_id": "my-profile-uuid",
        "text": "Hello, this is a programmatic voice sample.",
        "language": "en",
        "engine": "qwen",
        "normalize": true
      }' | jq .

Response:

{
  "id": "9f7c2a12-...",
  "profile_id": "my-profile-uuid",
  "text": "Hello, this is a programmatic voice sample.",
  "language": "en",
  "audio_path": "file:///voicebox/generated/9f7c2a12.wav",
  "duration": 2.34,
  "engine": "qwen",
  "status": "generating",
  "created_at": "2026-04-14T10:12:03.456Z"
}

Monitoring Generation Progress in Real Time

Voicebox offers two methods to track job status: Server-Sent Events (SSE) for push updates or simple HTTP polling.

Using Server-Sent Events

Connect to GET /generate/{id}/status using an EventSource. The endpoint defined in backend/routes/generations.py yields JSON payloads whenever backend/services/history.py updates the generation record.

const genId = '9f7c2a12-...';
const evtSource = new EventSource(`http://localhost:17493/generate/${genId}/status`);

evtSource.onmessage = (e) => {
  const payload = JSON.parse(e.data);
  console.log('Status update:', payload.status);
  if (payload.status === 'completed') {
    evtSource.close();
    // Fetch the audio file via payload.audio_path
  }
};

Polling with Python

If SSE is unavailable in your environment, poll the same endpoint with standard GET requests. The status progresses through "loading_model" → "generating" → "completed" or "failed".

import time, requests

base = "http://localhost:17493"
gen_id = "9f7c2a12-..."

while True:
    r = requests.get(f"{base}/generate/{gen_id}/status")
    data = r.json()
    print(f"Status: {data['status']}, duration: {data.get('duration')}")
    if data["status"] in ("completed", "failed"):
        break
    time.sleep(1)

Streaming and Downloading Audio Files

Once the status reaches "completed", the audio_path field contains a file:// URI to the WAV file. For applications that cannot wait for disk persistence, Voicebox provides a blocking POST /generate/stream endpoint that returns the raw audio bytes immediately after synthesis finishes.

const axios = require('axios');
const fs = require('fs');

async function streamWav() {
  const resp = await axios.post(
    'http://localhost:17493/generate/stream',
    {
      profile_id: 'my-profile-uuid',
      text: 'Streaming example.',
      language: 'en'
    },
    {
      responseType: 'stream',
      headers: { 'Content-Type': 'application/json' }
    }
  );

  const out = fs.createWriteStream('output.wav');
  resp.data.pipe(out);
  out.on('finish', () => console.log('WAV saved as output.wav'));
}

streamWav();

Retrying and Regenerating Audio

Voicebox supports non-destructive iteration through retry and regenerate endpoints.

Retry (POST /generate/{id}/retry): Re-runs the exact same text and seed, useful for transient GPU failures. This skips the effects pipeline to save processing time.

curl -X POST http://localhost:17493/generate/9f7c2a12-.../retry

Regenerate (POST /generate/{id}/regenerate): Creates a new variation with a random seed, stored as a separate version file (generationId_<random>.wav) labeled take-N. This allows A/B testing of voice renderings without losing previous outputs.

curl -X POST http://localhost:17493/generate/9f7c2a12-.../regenerate

Version tracking is managed through the SQLAlchemy models in backend/database/models.py.

The Generation Pipeline Deep Dive

Understanding the internal flow helps debug failures and optimize request patterns:

  1. Endpoint validation: backend/routes/generations.py validates the GenerationRequest against Pydantic models in backend/models.py, fetches the profile, and creates a database record via history.create_generation.

  2. Task queuing: task_manager.start_generation in backend/services/task_queue.py registers the job ID and hands the coroutine to enqueue_generation, which serializes execution to prevent GPU contention.

  3. Core synthesis: run_generation in backend/services/generation.py orchestrates the pipeline:

    • Loads the TTS engine via load_engine_model
    • Builds a voice prompt using profiles.create_voice_prompt_for_profile
    • Processes text through generate_chunked in backend/utils/chunked_tts.py, which splits long scripts, applies cross-fading, and optionally trims silence
    • Applies audio effects if specified using the Pedalboard-based pipeline in backend/utils/effects.py
    • Normalizes audio amplitude when normalize=True
    • Persists the final WAV and creates version records via _save_generate
  4. Status broadcasting: backend/services/history.py updates the generation status at each phase, which the SSE endpoint reads to push updates to clients.

Summary

  • Voicebox runs a FastAPI server on port 17493 with endpoints for generating, monitoring, and streaming speech
  • Submit text via POST /generate with a profile ID and receive a generation ID immediately while processing queues asynchronously
  • Track progress via Server-Sent Events at /generate/{id}/status or poll the SQLite-backed endpoint for status values: loading_model, generating, completed, or failed
  • Download completed WAV files from the returned audio_path or stream bytes directly via POST /generate/stream
  • Retry failed jobs with the same parameters or create variation takes using the retry and regenerate endpoints
  • All operations are backed by SQLite persistence defined in backend/database/models.py and an async task queue in backend/services/task_queue.py that serializes GPU-intensive work

Frequently Asked Questions

How do I authenticate requests to the Voicebox API?

The open-source jamiepine/voicebox implementation does not include authentication middleware; it assumes local network security. When exposing the server externally, place it behind a reverse proxy (such as Nginx or Traefik) with API key validation or OAuth2.

What is the maximum text length supported by the API?

The API accepts arbitrary text lengths in the text field, but the background service automatically chunks long scripts using the utility in backend/utils/chunked_tts.py, applying cross-fades between segments to create seamless audio without requiring manual splitting by the client.

Can I apply real-time audio effects through the API?

Yes, the GenerationRequest accepts effect parameters that backend/utils/effects.py processes using Pedalboard after the raw audio is generated. Alternatively, generate a clean version first, then access the effects processing separately or manage multiple versions through the database-backed version system.

How does the task queue handle concurrent requests?

The backend/services/task_queue.py implements an in-process async queue that serializes GPU-intensive generation tasks, ensuring only one runs at a time while maintaining a registry of active tasks. Client requests return immediately with a generation ID, and the actual inference runs in the background, allowing the API to accept new requests while previous ones complete.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →