What Is the OpenAI-Compatible API in VoiceStudio?

The OpenAI-compatible API in VoiceStudio is a FastAPI router layer that exposes local speech synthesis, recognition, and translation services through OpenAI-standard REST endpoints, enabling third-party clients to consume on-device models without modifying existing code.

The VoiceStudio repository (debpalash/VoiceStudio) ships with a compatibility bridge that maps OpenAI’s public API contract to its internal speech pipeline. This layer allows developers to point standard OpenAI clients at a local VoiceStudio instance—typically running on http://localhost:8000—to leverage high-performance, locally-hosted text-to-speech and speech-to-text models while maintaining full data privacy and avoiding external cloud calls.

Technical Architecture and Router Location

The compatibility layer is implemented as a FastAPI router located at api/routers/openai_compat.py. This module intercepts HTTP requests targeting OpenAI-style paths—such as /v1/audio/speech and /v1/audio/transcriptions—and dispatches them to VoiceStudio’s internal task workers for TTS, ASR, translation, and dubbing operations.

Architecturally, the router acts as a facade. Incoming requests are validated against Pydantic models mirroring OpenAI’s JSON schema, then routed to the appropriate backend engine through VoiceStudio’s internal transport layer rather than proxying to external cloud APIs.

Supported Endpoints and Operations

VoiceStudio’s OpenAI-compatible API mirrors the public /v1/audio/* namespace while substituting proprietary backend logic:

  • Text-to-Speech (TTS) — POST /v1/audio/speech accepts model, input, and voice parameters, returning raw audio bytes (WAV/MP3) generated by local neural vocoders.
  • Speech-to-Text (STT) — POST /v1/audio/transcriptions handles multipart file uploads, routing audio to local Whisper-style models (e.g., whisper-large-v3) without cloud API calls.
  • Translation & Dubbing — Extended routes support real-time translation workflows through the same REST contract, dispatching to VoiceStudio’s dub and asr task queues.

Integration Examples

Developers can consume VoiceStudio’s OpenAI-compatible layer via direct HTTP calls or official SDKs configured with a custom base URL.

Direct HTTP Requests

For text-to-speech synthesis:

import requests

base_url = "http://localhost:8000/v1/audio/speech"

payload = {
    "model": "gpt-4o-mini",  # Registered VoiceStudio model identifier

    "input": "Hello, this is a test.",
    "voice": "default"
}
resp = requests.post(base_url, json=payload)
audio_bytes = resp.content  # Raw audio data

For speech-to-text transcription:

import requests

files = {"file": open("sample.wav", "rb")}
params = {"model": "whisper-large-v3"}
resp = requests.post(
    "http://localhost:8000/v1/audio/transcriptions",
    data=params,
    files=files
)
print(resp.json()["text"])

Using the OpenAI Python SDK

Because the router preserves OpenAI’s request/response contract, you can override the base_url in the official client:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1")

result = client.audio.transcriptions.create(
    model="whisper-large-v3",
    file=open("sample.wav", "rb")
)
print(result.text)

Backend Dispatch and Safety Boundaries

The openai_compat router does not forward requests to OpenAI’s cloud servers. Instead, it validates incoming payloads and dispatches to VoiceStudio’s internal workers (tts, asr, dub) through the same transport layer used by the native API.

This architecture ensures that:

  • No external network calls are initiated for standard operations
  • The system gracefully rejects unimplemented OpenAI features (e.g., image generation) with explicit error responses
  • Audio watermarking and safety checks remain active regardless of the entry point

Testing and Validation

The repository includes comprehensive test coverage for the compatibility layer:

Summary

  • The OpenAI-compatible API in VoiceStudio is implemented as a FastAPI router at api/routers/openai_compat.py that intercepts OpenAI-standard REST calls
  • It exposes local TTS, STT, translation, and dubbing services through familiar /v1/audio/* endpoints without requiring cloud connectivity
  • Third-party integration is seamless: existing OpenAI client libraries work by changing only the base_url to the local VoiceStudio server
  • Safety and compliance are maintained through internal dispatch to VoiceStudio’s task workers, preserving watermarking and error-handling logic validated by the test suite

Frequently Asked Questions

What endpoints does the VoiceStudio OpenAI-compatible API support?

The router exposes endpoints under /v1/audio/ including speech synthesis (/speech), transcriptions (/transcriptions), and translations. These endpoints accept the same JSON schema and multipart forms as OpenAI’s official API, but execute against VoiceStudio’s local model registry rather than external cloud services.

Do I need to install the official OpenAI Python package to use VoiceStudio?

No. According to tests/test_llm_backend_not_blocked_by_missing_openai_package.py, the api/routers/openai_compat.py implementation has zero runtime dependency on the external openai package. You can use standard HTTP clients or the SDK if desired, but the router functions independently.

How does VoiceStudio handle audio watermarking through the OpenAI-compatible layer?

All audio processed through the compatibility layer undergoes the same watermarking pipeline as native API requests. The test suite in tests/test_watermark_route_coverage.py verifies that synthesized audio retains embedded provenance markers regardless of whether the request originates from the OpenAI-compatible router or VoiceStudio’s native frontend.

Can I extend the OpenAI-compatible API with custom VoiceStudio features?

Yes. Because the compatibility layer is a standard FastAPI router, developers can mount additional endpoints or middleware in api/routers/openai_compat.py. Custom endpoints can follow the OpenAI JSON schema for consistency while dispatching to VoiceStudio-specific services like advanced dubbing workflows or voice cloning pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →