VoiceStudio API Endpoints: Complete FastAPI Backend Reference

The VoiceStudio backend exposes over 50 RESTful endpoints organized into specialized FastAPI routers covering engine management, audio generation, speech processing, batch jobs, and OpenAI-compatible interfaces.

VoiceStudio is an open-source voice synthesis and processing platform built on FastAPI. This comprehensive guide maps every available VoiceStudio API endpoint across the backend router modules, providing exact file paths, HTTP methods, and integration patterns for building applications with the speech pipeline.

Engine Management and Health Checks

The Engine Management routes in backend/api/routers/engines.py serve as the central hub for discovering and configuring speech backends. These endpoints handle TTS (Text-to-Speech), ASR (Automatic Speech Recognition), LLM, and translation engines.

Key endpoints include:

  • List engines: GET /engines returns all available engine families, while GET /engines/tts, GET /engines/asr, and GET /engines/llm filter by category.
  • Select backend: POST /engines/select persists a backend choice for subsequent operations.
  • Health monitoring: GET /engines/{engine_id}/health performs a liveness check on a specific backend instance.
  • Translation management: GET /engines/translation lists available translation engines, with POST /engines/translation/{engine_id}/install and DELETE /engines/translation/{engine_id} managing their lifecycle.
import requests

# Check TTS engine availability

response = requests.get("http://localhost:8000/engines/tts")
available_engines = response.json()

# Verify engine health

health = requests.get(f"http://localhost:8000/engines/{engine_id}/health")
print(health.json()["status"])

Worker Orchestration

In backend/api/routers/workers.py, the Worker Control endpoints manage the distributed processing pool. Workers handle compute-intensive tasks like model inference and audio rendering.

  • GET /workers – Lists active workers and their current load.
  • POST /workers – Spawns a new worker process.
  • GET /workers/{worker_id} – Retrieves detailed status for a specific worker.
  • POST /workers/{worker_id}/join – Joins a worker to the control plane.
  • POST /workers/{worker_id}/shutdown – Gracefully stops a worker.

These routes include dependencies like Depends(require_admin) to protect administrative actions.


# Spawn a new worker

new_worker = requests.post("http://localhost:8000/workers")
worker_id = new_worker.json()["worker_id"]

# Shutdown worker when complete

requests.post(f"http://localhost:8000/workers/{worker_id}/shutdown")

Audio Generation and Streaming

The Audio Generation router in backend/api/routers/generation.py handles asynchronous TTS job submission, while backend/api/routers/tts_stream.py provides Server-Sent Events (SSE) for real-time audio delivery.

Generation endpoints:

  • POST /generation – Submits a TTS generation job and returns a job_id.
  • GET /generation/{job_id}/status – Polls job progress.
  • GET /generation/{job_id}/result – Retrieves the final audio file.
  • POST /generation/{job_id}/cancel – Aborts a running job.

Streaming endpoints:

  • GET /tts_stream/{task_id} – Opens an SSE connection that streams progressive audio chunks as they are generated.
import requests

# Submit generation job

job = requests.post("http://localhost:8000/generation", json={
    "text": "Hello world",
    "voice_id": "en-US-Aria"
})
job_id = job.json()["job_id"]

# Stream chunks via SSE

url = f"http://localhost:8000/tts_stream/{job_id}"
with requests.get(url, stream=True) as r:
    for line in r.iter_lines():
        if line:
            process_audio_chunk(line)

Speech-to-Text and Text-to-Speech

The Speech Platform router in backend/api/routers/speech_platform.py provides a unified service layer for synchronous speech operations:

  • POST /speech_platform/tts – Synthesizes speech immediately (blocking).
  • POST /speech_platform/stt – Transcribes uploaded audio.

For Capture and Acoustic Echo Cancellation (AEC), backend/api/routers/capture.py and backend/api/routers/capture_ws.py offer:

  • POST /capture – Records audio capture.
  • GET /capture_ws – WebSocket endpoint for live audio streaming with AEC processing.

# Synchronous TTS

response = requests.post("http://localhost:8000/speech_platform/tts", 
    json={"text": "Processing complete"})
audio_data = response.content

OpenAI Compatibility Layer

VoiceStudio provides drop-in compatibility with OpenAI clients through backend/api/routers/openai_compat.py. This router maps standard OpenAI audio endpoints to VoiceStudio's internal pipeline:

  • POST /v1/audio/tts – OpenAI-compatible TTS endpoint accepting model, voice, and input parameters.
  • POST /v1/audio/transcriptions – OpenAI-compatible STT endpoint for transcription requests.

This allows existing applications using OpenAI SDKs to switch to VoiceStudio by changing only the base URL.

Dubbing, Audiobook, and Batch Processing

For complex production workflows, VoiceStudio organizes functionality across multiple specialized routers.

Dub & Export (backend/api/routers/dub_core.py, dub_generate.py, dub_export.py):

  • POST /dub_core/submit – Submits a dubbing request with source and target languages.
  • GET /dub_generate/{task_id}/status – Streams generation status.
  • GET /dub_export/{task_id}/file – Downloads the exported dubbed file.

Long-Form Content (backend/api/routers/audiobook.py, longform_jobs.py):

  • POST /audiobook – Creates an audiobook production job.
  • GET /audiobook/{task_id}/stream – SSE stream of completed chapters.

Batch Processing (backend/api/routers/batch.py):

  • POST /batch – Submits multiple generation jobs as a single batch.
  • GET /batch/{batch_id}/status – Monitors overall batch progress.

System Administration and Settings

Administrative endpoints reside in backend/api/routers/system.py and backend/api/routers/settings.py:

  • GET /system/info – Returns runtime information.
  • POST /system/restart – Triggers process restart (admin-only).
  • GET /settings – Retrieves user preferences.
  • POST /settings – Updates configuration.

Media Tools (backend/api/routers/media_tools.py) provide utility functions:

  • POST /media_tools/convert – Converts between audio/video formats.
  • GET /media_tools/metadata – Probes file metadata.

Authentication, Profiles, and Utilities

Security routes in backend/api/routers/auth.py handle session management:

  • POST /auth/login – Authenticates users.
  • POST /auth/logout – Terminates sessions.

The Profiles router (backend/api/routers/profiles.py) manages user-specific configurations:

  • GET /profiles – Lists profiles.
  • POST /profiles – Creates or modifies a profile.

Additional utility endpoints include:

Summary

  • VoiceStudio organizes its FastAPI backend into modular routers under backend/api/routers/, each handling specific functional domains.
  • Engine management endpoints in engines.py control TTS, ASR, LLM, and translation backends with health monitoring.
  • Audio generation supports both synchronous (/generation) and streaming (/tts_stream) patterns via SSE.
  • OpenAI compatibility is provided through openai_compat.py, enabling drop-in replacement for OpenAI audio APIs.
  • Worker orchestration allows dynamic scaling of inference workers through the workers.py routes.
  • Dubbing, audiobook production, and batch processing provide enterprise-grade workflow automation.

Frequently Asked Questions

How do I check if a specific TTS engine is healthy?

Send a GET request to /engines/{engine_id}/health as defined in backend/api/routers/engines.py. This endpoint performs a liveness check on the backend instance and returns a status object indicating whether the engine is ready to accept generation jobs.

What is the difference between /generation and /tts_stream endpoints?

The /generation endpoint in backend/api/routers/generation.py works asynchronously—you submit a job and poll for completion—while /tts_stream in backend/api/routers/tts_stream.py uses Server-Sent Events (SSE) to deliver audio chunks progressively as they are synthesized, enabling real-time playback.

Does VoiceStudio support OpenAI-compatible API clients?

Yes. The backend/api/routers/openai_compat.py module exposes POST /v1/audio/tts and POST /v1/audio/transcriptions, which accept the same request schemas as OpenAI's audio endpoints. You can use VoiceStudio as a backend for existing OpenAI SDK implementations by pointing the base URL to your VoiceStudio instance.

How are batch jobs submitted and monitored?

Submit a batch by sending a POST request to /batch with a list of generation tasks. The endpoint returns a batch_id that you can use with GET /batch/{batch_id}/status to track overall progress. This functionality is implemented in backend/api/routers/batch.py and supports parallel processing of multiple audio generation tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →