How to Select a TTS Engine in VoiceStudio Programmatically

You can select a TTS engine in VoiceStudio by sending the engine key in your JSON payload to the /v1/generation endpoint, or by setting the OMNIVOICE_TTS_BACKEND environment variable before launching the server.

VoiceStudio, from the debpalash/VoiceStudio repository, exposes all backend text-to-speech engines through a central registry. You can select a TTS engine in VoiceStudio programmatically on every request, which makes it straightforward to script batch jobs or A/B test different synthesizers without touching the web UI.

How the TTS Engine Registry Works

The backend maintains a hardcoded registry of supported synthesizers that the API uses for discovery and dispatch.

Central Engine Registry

In backend/services/tts_backend.py, VoiceStudio defines a _REGISTRY dictionary that maps unique engine IDs to concrete backend classes. The list_backends() helper returns the currently available providers. Supported IDs include indextts, omnivoice, and voxcpm2.

Discovery via the /engines Endpoint

The API router in backend/api/routers/engines.py reads that same registry and exposes an /engines endpoint. The UI and external clients query this endpoint to discover which engines are installed and ready for synthesis.

Sending the Engine Field in an API Request

When you POST to the generation handler, VoiceStudio inspects the incoming JSON for an engine key and routes the job accordingly.

REST API Payload Structure

As implemented in backend/api/generation.py, the handler extracts the engine value from the request body and dispatches the task to the matching backend. If you omit the key, the system falls back to the default configured in the Settings UI or the OMNIVOICE_TTS_BACKEND environment variable.

A valid payload includes the following fields:

  • engine — the backend ID (e.g., "indextts")
  • operation — typically "tts"
  • params — engine-specific arguments such as {"text": "Hello, world!"}
  • deadline_seconds — timeout for the synthesis job

Python requests Example

import os
import requests

# Optional: override the default engine for the whole process

os.environ["OMNIVOICE_TTS_BACKEND"] = "indextts"

url = "http://localhost:8000/v1/generation"
payload = {
    "engine": "indextts",          # <-- select the engine here

    "operation": "tts",
    "params": {"text": "Hello, world!"},
    "deadline_seconds": 60,
}
resp = requests.post(url, json=payload)
audio_bytes = resp.content          # binary WAV/MP3 data

with open("out.wav", "wb") as f:
    f.write(audio_bytes)

curl Example

curl -X POST http://localhost:8000/v1/generation \
     -H "Content-Type: application/json" \
     -d '{
           "engine": "omnivoice",
           "operation": "tts",
           "params": {"text": "Voice Studio demo"},
           "deadline_seconds": 30
         }' --output demo.wav

Using the Python SDK

If you import the VoiceStudio client library, you can pass the engine directly to the synthesize() method:

from voice_studio import VoiceStudioClient

client = VoiceStudioClient(base_url="http://localhost:8000")
audio = client.synthesize(
    text="Programmatic selection test",
    engine="indextts",        # explicit engine selection

)
with open("tts.wav", "wb") as f:
    f.write(audio)

Configuring a Default Engine via Environment Variables

You can set a global default so that every request without an explicit engine field uses the same backend. Export the variable before launching the server or worker process:

export OMNIVOICE_TTS_BACKEND=voxcpm2   # will be used unless overridden per‑request

voice-studio   # start the server / UI

This value acts as the fallback when the JSON payload does not specify an engine.

Where Engine Resolution Happens in the Source Code

Understanding the dispatch path helps when debugging or extending VoiceStudio:

Summary

  • Pass the engine key in your JSON payload to POST /v1/generation to choose a synthesizer per-request.
  • Valid engine IDs are defined in backend/services/tts_backend.py and exposed through the /engines endpoint.
  • Omitting the engine field falls back to the OMNIVOICE_TTS_BACKEND environment variable or the Settings UI default.
  • The dispatch logic lives in backend/api/generation.py, which matches the request engine to the registry and runs synthesis.

Frequently Asked Questions

What happens if I omit the engine field in the generation payload?

VoiceStudio falls back to the default engine configured in the Settings UI or the OMNIVOICE_TTS_BACKEND environment variable. The handler in backend/api/generation.py performs this resolution before dispatching the job.

How do I list all available TTS engines programmatically?

Query the /engines endpoint or import the registry helper from backend/services/tts_backend.py. The list_backends() function returns the installed engine IDs and metadata that the UI uses to populate its engine picker.

Can I switch engines without restarting the VoiceStudio server?

Yes. Because the engine field is read from each incoming JSON payload at request time, you can send one request with "engine": "indextts" and the next with "engine": "omnivoice" without restarting the process. Only the global default requires an application restart to change via environment variables.

Where does VoiceStudio validate the engine ID?

Validation occurs in the generation handler inside backend/api/generation.py. The handler resolves the provided string against the _REGISTRY dictionary defined in backend/services/tts_backend.py. If the ID is missing, the system uses the configured fallback instead of raising an error.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →