How to Use the OpenAI‑Compatible API in VoiceStudio: Complete Implementation Guide

VoiceStudio provides a drop‑in replacement for OpenAI’s audio endpoints at the /v1/audio prefix, allowing you to use standard OpenAI client libraries with VoiceStudio’s TTS and STT backends.

The OpenAI‑compatible API in VoiceStudio is implemented as a FastAPI router that mirrors the official OpenAI specification while adding support for VoiceStudio‑specific engines and voice profiles. Located in backend/api/routers/openai_compat.py, this compatibility layer enables seamless integration with existing applications that expect OpenAI‑style request and response formats.

Available Endpoints and Routes

VoiceStudio exposes three primary routes under the /v1/audio path, each implemented as a distinct function in the router:

  • /v1/audio/speech – The text‑to‑speech (TTS) endpoint handled by create_speech (lines 21‑30). Accepts a JSON payload describing the text, voice, and model, returning synthesized audio.
  • /v1/audio/transcriptions – The speech‑to‑text (STT) endpoint handled by create_transcription (lines 11‑13). Accepts multipart form data with an audio file and returns transcription text.
  • /v1/audio/voices – The voice listing endpoint handled by list_voices (lines 79‑87). Returns available voices and supported engine IDs.

Each route maintains compatibility with OpenAI’s REST contract while supporting VoiceStudio’s extended parameter sets.

Request and Response Schemas

The router defines Pydantic models that extend OpenAI’s specifications with VoiceStudio‑specific fields.

SpeechRequest

The SpeechRequest schema (defined at lines 45‑78) accepts standard fields like model, input, voice, and response_format, plus VoiceStudio extensions:

  • language – Target language code for synthesis
  • duration – Desired output duration (for specific engines)
  • seed – Reproducibility seed for deterministic generation

Transcription Responses

For STT operations, the API returns two possible schemas:

  • TranscriptionResponse (lines 33‑37) – A simple JSON object containing only the text field.
  • VerboseTranscriptionResponse (lines 39‑46) – Returned when response_format=verbose_json, containing additional metadata like duration and language detection.

Engine Resolution and Model Mapping

VoiceStudio maps the OpenAI model identifiers to internal TTS engines through the _resolve_engine helper function (lines 61‑73).

Model alias resolution works as follows:

  • tts-1 and tts-1-hd automatically resolve to the currently active engine configured in VoiceStudio.
  • Custom identifiers (e.g., voxcpm2, cosyvoice) must match a registered backend ID in the system.

This resolution occurs before synthesis begins, ensuring the request routes to the correct GPU-backed service.

Admission Controls and Error Handling

Before processing TTS requests, create_speech implements admission controls (lines 26‑34) to prevent resource exhaustion:

  • HTTP 400 – Returned when the requested engine is unavailable or the specified model cannot be resolved.
  • HTTP 429 – Returned when the GPU pool is saturated and cannot accept new synthesis jobs.

These checks ensure VoiceStudio maintains quality of service under high load.

Supported Audio Formats

VoiceStudio supports multiple output formats through the response_format parameter, handled by the _encode_audio utility (lines 5‑63).

TTS formats: mp3, wav, opus, flac, aac, pcm

STT response formats: json, text, verbose_json, srt, vtt

The router automatically sets the appropriate MIME type and file extension based on your selection.

Implementation Examples

Text‑to‑Speech with cURL

Send a POST request to generate audio:

curl -X POST https://localhost:3900/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
        "model": "tts-1",
        "input": "Hello, VoiceStudio!",
        "voice": "default",
        "response_format": "mp3"
      }' --output hello.mp3

Replace https://localhost:3900 with your VoiceStudio FastAPI server address.

Using the Official OpenAI Python Client

The openai package (declared in pyproject.toml) works transparently with VoiceStudio by pointing the base_url to your local instance:

import openai

client = openai.OpenAI(
    base_url="http://localhost:3900/v1",  # VoiceStudio endpoint

    api_key="any-string"                  # No authentication required locally

)

# Generate speech

speech = client.audio.speech.create(
    model="tts-1",
    input="Welcome to the OpenAI‑compatible API!",
    voice="default",
    response_format="mp3"
)
with open("welcome.mp3", "wb") as f:
    f.write(speech.read())

# Transcribe audio

with open("sample.wav", "rb") as audio_file:
    transcription = client.audio.transcriptions.create(
        model="whisper-1",
        file=audio_file,
        response_format="text"
    )
print(transcription.text)

Listing Available Voices

Query the voices endpoint to discover available profiles and engine mappings:

curl http://localhost:3900/v1/audio/voices

Example response:

{
  "voices": [
    {"voice_id":"alloy","name":"Alloy","type":"openai_alias","description":"OpenAI 'alloy' voice — maps to the active VoiceStudio engine's default voice."},
    {"voice_id":"my-profile","name":"My Voice","type":"profile","language":"en"}
  ],
  "engines": ["omnivice", "voxcpm2", "cosyvoice"]
}

Key Source Files

Understanding the architecture requires familiarity with these modules:

Summary

  • VoiceStudio’s OpenAI‑compatible API exposes /v1/audio/speech, /v1/audio/transcriptions, and /v1/audio/voices endpoints that mirror OpenAI’s specification.
  • The SpeechRequest schema extends OpenAI’s format with VoiceStudio‑specific fields like duration and seed.
  • Model aliases tts-1 and tts-1-hd resolve to the active engine, while custom IDs map directly to registered backends like voxcpm2.
  • The system returns HTTP 400 for invalid engines and HTTP 429 when GPU resources are exhausted.
  • You can use standard OpenAI client libraries by setting base_url to your VoiceStudio instance address.

Frequently Asked Questions

How do I switch between different TTS engines in VoiceStudio?

Pass the specific engine ID as the model parameter in your request. While tts-1 uses the default active engine, you can specify registered backends like voxcpm2 or cosyvoice directly. The _resolve_engine function in backend/api/routers/openai_compat.py handles this mapping.

What audio file formats are supported for transcription?

The transcription endpoint accepts standard audio inputs and can return responses in json, text, verbose_json, srt, or vtt formats. Set the response_format parameter in your request to control the output structure.

Do I need an API key to use the OpenAI‑compatible endpoints?

For local deployments, no authentication is required. You can pass any string as the api_key parameter when initializing the OpenAI client. Production deployments may implement additional security layers outside the compatibility router.

Can I use VoiceStudio’s extended parameters with standard OpenAI clients?

Yes. The SpeechRequest schema accepts additional fields like language, duration, and seed that are specific to VoiceStudio. Standard OpenAI clients will pass these through in the JSON payload, and VoiceStudio will process them while maintaining compatibility with the base specification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →