Using Voicebox REST API for Programmatic Voice Generation: A Complete Integration Guide
Voicebox exposes a FastAPI-based REST interface on port 17493 that queues text-to-speech requests via POST /generate, streams real-time progress through Server-Sent Events at /generate/{id}/status, and delivers WAV audio files, enabling fully automated voice synthesis without the desktop UI.
Voicebox is an open-source text-to-speech application that provides a comprehensive REST API for programmatic voice generation. According to the jamiepine/voicebox source code, the FastAPI server handles asynchronous GPU task queuing, real-time status monitoring, and audio streaming through well-defined endpoints. This guide covers the exact API contract, request lifecycle, and integration patterns needed to embed Voicebox into automated workflows.
Understanding the API Architecture
The Voicebox server is built on FastAPI with endpoints defined in backend/routes/generations.py. All data validation uses Pydantic models located in backend/models.py, including GenerationRequest and GenerationResponse.
Long-running TTS jobs are not executed synchronously. Instead, backend/services/task_queue.py maintains an in-process async queue that serializes GPU work, ensuring only one generation runs at a time while tracking active tasks. When a job completes, backend/services/history.py updates the SQLite database (defined in backend/database/models.py) and pushes status changes to connected SSE clients.
Submitting Your First Generation Request
To initiate speech synthesis, send a POST request to /generate with a JSON payload matching the GenerationRequest schema. Required fields include profile_id (UUID of the voice profile), text (the script), and optionally language, engine (defaults to the profile’s setting or "qwen"), and normalize (boolean for audio normalization).
The endpoint immediately returns a GenerationResponse containing the generation id, initial status set to "generating", and the target audio_path. The actual inference runs in the background via task_manager.start_generation and enqueue_generation as implemented in the task queue service.
curl -X POST http://localhost:17493/generate \
-H "Content-Type: application/json" \
-d '{
"profile_id": "my-profile-uuid",
"text": "Hello, this is a programmatic voice sample.",
"language": "en",
"engine": "qwen",
"normalize": true
}' | jq .
Response:
{
"id": "9f7c2a12-...",
"profile_id": "my-profile-uuid",
"text": "Hello, this is a programmatic voice sample.",
"language": "en",
"audio_path": "file:///voicebox/generated/9f7c2a12.wav",
"duration": 2.34,
"engine": "qwen",
"status": "generating",
"created_at": "2026-04-14T10:12:03.456Z"
}
Monitoring Generation Progress in Real Time
Voicebox offers two methods to track job status: Server-Sent Events (SSE) for push updates or simple HTTP polling.
Using Server-Sent Events
Connect to GET /generate/{id}/status using an EventSource. The endpoint defined in backend/routes/generations.py yields JSON payloads whenever backend/services/history.py updates the generation record.
const genId = '9f7c2a12-...';
const evtSource = new EventSource(`http://localhost:17493/generate/${genId}/status`);
evtSource.onmessage = (e) => {
const payload = JSON.parse(e.data);
console.log('Status update:', payload.status);
if (payload.status === 'completed') {
evtSource.close();
// Fetch the audio file via payload.audio_path
}
};
Polling with Python
If SSE is unavailable in your environment, poll the same endpoint with standard GET requests. The status progresses through "loading_model" → "generating" → "completed" or "failed".
import time, requests
base = "http://localhost:17493"
gen_id = "9f7c2a12-..."
while True:
r = requests.get(f"{base}/generate/{gen_id}/status")
data = r.json()
print(f"Status: {data['status']}, duration: {data.get('duration')}")
if data["status"] in ("completed", "failed"):
break
time.sleep(1)
Streaming and Downloading Audio Files
Once the status reaches "completed", the audio_path field contains a file:// URI to the WAV file. For applications that cannot wait for disk persistence, Voicebox provides a blocking POST /generate/stream endpoint that returns the raw audio bytes immediately after synthesis finishes.
const axios = require('axios');
const fs = require('fs');
async function streamWav() {
const resp = await axios.post(
'http://localhost:17493/generate/stream',
{
profile_id: 'my-profile-uuid',
text: 'Streaming example.',
language: 'en'
},
{
responseType: 'stream',
headers: { 'Content-Type': 'application/json' }
}
);
const out = fs.createWriteStream('output.wav');
resp.data.pipe(out);
out.on('finish', () => console.log('WAV saved as output.wav'));
}
streamWav();
Retrying and Regenerating Audio
Voicebox supports non-destructive iteration through retry and regenerate endpoints.
Retry (POST /generate/{id}/retry): Re-runs the exact same text and seed, useful for transient GPU failures. This skips the effects pipeline to save processing time.
curl -X POST http://localhost:17493/generate/9f7c2a12-.../retry
Regenerate (POST /generate/{id}/regenerate): Creates a new variation with a random seed, stored as a separate version file (generationId_<random>.wav) labeled take-N. This allows A/B testing of voice renderings without losing previous outputs.
curl -X POST http://localhost:17493/generate/9f7c2a12-.../regenerate
Version tracking is managed through the SQLAlchemy models in backend/database/models.py.
The Generation Pipeline Deep Dive
Understanding the internal flow helps debug failures and optimize request patterns:
-
Endpoint validation:
backend/routes/generations.pyvalidates theGenerationRequestagainst Pydantic models inbackend/models.py, fetches the profile, and creates a database record viahistory.create_generation. -
Task queuing:
task_manager.start_generationinbackend/services/task_queue.pyregisters the job ID and hands the coroutine toenqueue_generation, which serializes execution to prevent GPU contention. -
Core synthesis:
run_generationinbackend/services/generation.pyorchestrates the pipeline:- Loads the TTS engine via
load_engine_model - Builds a voice prompt using
profiles.create_voice_prompt_for_profile - Processes text through
generate_chunkedinbackend/utils/chunked_tts.py, which splits long scripts, applies cross-fading, and optionally trims silence - Applies audio effects if specified using the Pedalboard-based pipeline in
backend/utils/effects.py - Normalizes audio amplitude when
normalize=True - Persists the final WAV and creates version records via
_save_generate
- Loads the TTS engine via
-
Status broadcasting:
backend/services/history.pyupdates the generation status at each phase, which the SSE endpoint reads to push updates to clients.
Summary
- Voicebox runs a FastAPI server on port
17493with endpoints for generating, monitoring, and streaming speech - Submit text via
POST /generatewith a profile ID and receive a generation ID immediately while processing queues asynchronously - Track progress via Server-Sent Events at
/generate/{id}/statusor poll the SQLite-backed endpoint for status values:loading_model,generating,completed, orfailed - Download completed WAV files from the returned
audio_pathor stream bytes directly viaPOST /generate/stream - Retry failed jobs with the same parameters or create variation takes using the retry and regenerate endpoints
- All operations are backed by SQLite persistence defined in
backend/database/models.pyand an async task queue inbackend/services/task_queue.pythat serializes GPU-intensive work
Frequently Asked Questions
How do I authenticate requests to the Voicebox API?
The open-source jamiepine/voicebox implementation does not include authentication middleware; it assumes local network security. When exposing the server externally, place it behind a reverse proxy (such as Nginx or Traefik) with API key validation or OAuth2.
What is the maximum text length supported by the API?
The API accepts arbitrary text lengths in the text field, but the background service automatically chunks long scripts using the utility in backend/utils/chunked_tts.py, applying cross-fades between segments to create seamless audio without requiring manual splitting by the client.
Can I apply real-time audio effects through the API?
Yes, the GenerationRequest accepts effect parameters that backend/utils/effects.py processes using Pedalboard after the raw audio is generated. Alternatively, generate a clean version first, then access the effects processing separately or manage multiple versions through the database-backed version system.
How does the task queue handle concurrent requests?
The backend/services/task_queue.py implements an in-process async queue that serializes GPU-intensive generation tasks, ensuring only one runs at a time while maintaining a registry of active tasks. Client requests return immediately with a generation ID, and the actual inference runs in the background, allowing the API to accept new requests while previous ones complete.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →