# How to Use the OpenAI‑Compatible API in VoiceStudio: Complete Implementation Guide

> Integrate the OpenAI-compatible API into VoiceStudio effortlessly. Use standard OpenAI clients with VoiceStudio's TTS and STT backends for seamless audio processing.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-12

---

**VoiceStudio provides a drop‑in replacement for OpenAI’s audio endpoints at the `/v1/audio` prefix, allowing you to use standard OpenAI client libraries with VoiceStudio’s TTS and STT backends.**

The OpenAI‑compatible API in VoiceStudio is implemented as a FastAPI router that mirrors the official OpenAI specification while adding support for VoiceStudio‑specific engines and voice profiles. Located in [`backend/api/routers/openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/openai_compat.py), this compatibility layer enables seamless integration with existing applications that expect OpenAI‑style request and response formats.

## Available Endpoints and Routes

VoiceStudio exposes three primary routes under the `/v1/audio` path, each implemented as a distinct function in the router:

- **`/v1/audio/speech`** – The **text‑to‑speech (TTS)** endpoint handled by `create_speech` (lines 21‑30). Accepts a JSON payload describing the text, voice, and model, returning synthesized audio.
- **`/v1/audio/transcriptions`** – The **speech‑to‑text (STT)** endpoint handled by `create_transcription` (lines 11‑13). Accepts multipart form data with an audio file and returns transcription text.
- **`/v1/audio/voices`** – The **voice listing** endpoint handled by `list_voices` (lines 79‑87). Returns available voices and supported engine IDs.

Each route maintains compatibility with OpenAI’s REST contract while supporting VoiceStudio’s extended parameter sets.

## Request and Response Schemas

The router defines Pydantic models that extend OpenAI’s specifications with VoiceStudio‑specific fields.

### SpeechRequest

The `SpeechRequest` schema (defined at lines 45‑78) accepts standard fields like `model`, `input`, `voice`, and `response_format`, plus VoiceStudio extensions:

- **`language`** – Target language code for synthesis
- **`duration`** – Desired output duration (for specific engines)
- **`seed`** – Reproducibility seed for deterministic generation

### Transcription Responses

For STT operations, the API returns two possible schemas:

- **`TranscriptionResponse`** (lines 33‑37) – A simple JSON object containing only the `text` field.
- **`VerboseTranscriptionResponse`** (lines 39‑46) – Returned when `response_format=verbose_json`, containing additional metadata like duration and language detection.

## Engine Resolution and Model Mapping

VoiceStudio maps the OpenAI model identifiers to internal TTS engines through the `_resolve_engine` helper function (lines 61‑73).

**Model alias resolution works as follows:**

- **`tts-1`** and **`tts-1-hd`** automatically resolve to the currently *active* engine configured in VoiceStudio.
- **Custom identifiers** (e.g., `voxcpm2`, `cosyvoice`) must match a registered backend ID in the system.

This resolution occurs before synthesis begins, ensuring the request routes to the correct GPU-backed service.

## Admission Controls and Error Handling

Before processing TTS requests, `create_speech` implements admission controls (lines 26‑34) to prevent resource exhaustion:

- **HTTP 400** – Returned when the requested engine is unavailable or the specified model cannot be resolved.
- **HTTP 429** – Returned when the GPU pool is saturated and cannot accept new synthesis jobs.

These checks ensure VoiceStudio maintains quality of service under high load.

## Supported Audio Formats

VoiceStudio supports multiple output formats through the `response_format` parameter, handled by the `_encode_audio` utility (lines 5‑63).

**TTS formats:** `mp3`, `wav`, `opus`, `flac`, `aac`, `pcm`

**STT response formats:** `json`, `text`, `verbose_json`, `srt`, `vtt`

The router automatically sets the appropriate MIME type and file extension based on your selection.

## Implementation Examples

### Text‑to‑Speech with cURL

Send a POST request to generate audio:

```bash
curl -X POST https://localhost:3900/v1/audio/speech \
  -H "Content-Type: application/json" \
  -d '{
        "model": "tts-1",
        "input": "Hello, VoiceStudio!",
        "voice": "default",
        "response_format": "mp3"
      }' --output hello.mp3

```

Replace `https://localhost:3900` with your VoiceStudio FastAPI server address.

### Using the Official OpenAI Python Client

The `openai` package (declared in [`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml)) works transparently with VoiceStudio by pointing the `base_url` to your local instance:

```python
import openai

client = openai.OpenAI(
    base_url="http://localhost:3900/v1",  # VoiceStudio endpoint

    api_key="any-string"                  # No authentication required locally

)

# Generate speech

speech = client.audio.speech.create(
    model="tts-1",
    input="Welcome to the OpenAI‑compatible API!",
    voice="default",
    response_format="mp3"
)
with open("welcome.mp3", "wb") as f:
    f.write(speech.read())

# Transcribe audio

with open("sample.wav", "rb") as audio_file:
    transcription = client.audio.transcriptions.create(
        model="whisper-1",
        file=audio_file,
        response_format="text"
    )
print(transcription.text)

```

### Listing Available Voices

Query the voices endpoint to discover available profiles and engine mappings:

```bash
curl http://localhost:3900/v1/audio/voices

```

**Example response:**

```json
{
  "voices": [
    {"voice_id":"alloy","name":"Alloy","type":"openai_alias","description":"OpenAI 'alloy' voice — maps to the active VoiceStudio engine's default voice."},
    {"voice_id":"my-profile","name":"My Voice","type":"profile","language":"en"}
  ],
  "engines": ["omnivice", "voxcpm2", "cosyvoice"]
}

```

## Key Source Files

Understanding the architecture requires familiarity with these modules:

- **[`backend/api/routers/openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/openai_compat.py)** – Implements the `/v1/audio/*` routes and request validation schemas.
- **[`pyproject.toml`](https://github.com/debpalash/VoiceStudio/blob/main/pyproject.toml)** – Declares the `openai` Python package dependency required for client usage.
- **[`services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/tts_backend.py)** – Provides the `_resolve_engine` logic and TTS generation implementation.
- **[`services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/asr_backend.py)** – Supplies ASR backend selection for the transcriptions endpoint.
- **[`services/audio_io.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/audio_io.py)** – Handles safe audio encoding across format types.
- **[`services/text_normalization.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/text_normalization.py)** – Normalizes input text before synthesis in `create_speech`.

## Summary

- VoiceStudio’s OpenAI‑compatible API exposes **`/v1/audio/speech`**, **`/v1/audio/transcriptions`**, and **`/v1/audio/voices`** endpoints that mirror OpenAI’s specification.
- The **SpeechRequest** schema extends OpenAI’s format with VoiceStudio‑specific fields like `duration` and `seed`.
- Model aliases **`tts-1`** and **`tts-1-hd`** resolve to the active engine, while custom IDs map directly to registered backends like `voxcpm2`.
- The system returns **HTTP 400** for invalid engines and **HTTP 429** when GPU resources are exhausted.
- You can use standard OpenAI client libraries by setting `base_url` to your VoiceStudio instance address.

## Frequently Asked Questions

### How do I switch between different TTS engines in VoiceStudio?

Pass the specific engine ID as the `model` parameter in your request. While `tts-1` uses the default active engine, you can specify registered backends like `voxcpm2` or `cosyvoice` directly. The `_resolve_engine` function in [`backend/api/routers/openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/openai_compat.py) handles this mapping.

### What audio file formats are supported for transcription?

The transcription endpoint accepts standard audio inputs and can return responses in `json`, `text`, `verbose_json`, `srt`, or `vtt` formats. Set the `response_format` parameter in your request to control the output structure.

### Do I need an API key to use the OpenAI‑compatible endpoints?

For local deployments, no authentication is required. You can pass any string as the `api_key` parameter when initializing the OpenAI client. Production deployments may implement additional security layers outside the compatibility router.

### Can I use VoiceStudio’s extended parameters with standard OpenAI clients?

Yes. The `SpeechRequest` schema accepts additional fields like `language`, `duration`, and `seed` that are specific to VoiceStudio. Standard OpenAI clients will pass these through in the JSON payload, and VoiceStudio will process them while maintaining compatibility with the base specification.