How the OpenAI-Compatible Audio API Maps `model` to Engine Selection and Supports Output Formats

The OpenAI-compatible audio router in VoiceStudio translates the model parameter into specific TTS or STT engine instances and supports six audio output formats with graceful fallback handling.

VoiceStudio provides a drop-in replacement for OpenAI's speech endpoints through its /v1/audio/* routes. This implementation, found in backend/api/routers/openai_compat.py, reuses the platform's native engine architecture while maintaining full compatibility with OpenAI's API contract. Understanding how the model parameter drives engine selection and which formats are supported helps developers integrate VoiceStudio seamlessly into existing workflows.

Model-to-Engine Mapping for Text-to-Speech

When clients POST to /v1/audio/speech, the model field undergoes resolution through the _resolve_engine helper function.

OpenAI Aliases vs. Direct Engine IDs

The router accepts two categories of model values:

Model value Mapping behavior
"tts-1" or "tts-1-hd" Returns the active TTS backend—whatever engine is currently selected by the user in the VoiceStudio UI
Valid engine IDs ("omnivoice", "voxcpm2", "cosyvoice", etc.) Looks up the backend class via services.tts_backend.get_backend_class and returns a singleton instance through get_engine_instance_for
Unknown or invalid value Raises 400 Bad Request with a list of supported engine IDs

The resolution logic spans lines 61–102 of openai_compat.py. For registered backends, the singleton pattern prevents repeated initialization overhead while still permitting hot-swapping of the active engine.

Model-to-Engine Mapping for Speech-to-Text

The STT endpoint /v1/audio/transcriptions applies parallel logic with Whisper-specific defaults.

The model form field defaults to "whisper-1", which maps to the active ASR backend. Any other string is interpreted as a specific VoiceStudio ASR engine identifier such as whisperx or faster-whisper (lines 15–20).

This design ensures that existing OpenAI SDK clients work without modification, while power users can target specific engines by passing their VoiceStudio IDs.

Supported Output Formats in TTS

The response_format field in SpeechRequest accepts a Literal type with six possible values:


"mp3", "opus", "aac", "flac", "wav", "pcm"

After synthesis produces a raw torch.Tensor, the _encode_audio function routes to format-specific encoders:

Format Encoding implementation
wav torchaudio.save(..., format="wav") → MIME audio/wav
flac torchaudio.save(..., format="flac") → MIME audio/flac
mp3 Attempts format="mp3"; falls back to WAV if ffmpeg is unavailable
opus Attempts format="ogg" (Opus container); falls back to WAV on failure
pcm Returns raw 16-bit little-endian PCM bytes without container
aac Not directly supported; automatically falls back to WAV

The encoder selection spans lines 205–262. Error handling is defensive: when preferred containers cannot be written, the router degrades to WAV rather than failing the request. The final StreamingResponse carries correct MIME types and file extensions to ensure proper client-side handling.

Architecture Benefits

Three design decisions make this implementation robust:

  • Engine-agnostic routing — The same admission-control logic from native endpoints enforces GPU/CPU caps and model-load budgets
  • Singleton backend instances — Cached per-engine classes eliminate repeated sidecar spawns
  • Graceful degradation — Missing ffmpeg or unsupported codecs trigger automatic fallback to universally compatible WAV

Code Examples

TTS with OpenAI Alias

import requests

resp = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={
        "model": "tts-1",          # maps to currently active backend

        "input": "Hello, world!",
        "voice": "default",
        "response_format": "mp3"
    },
)
with open("hello.mp3", "wb") as f:
    f.write(resp.content)

STT with Specific Engine

import requests

files = {"file": open("speech.wav", "rb")}
data = {
    "model": "faster-whisper",   # selects faster-whisper backend directly

    "response_format": "text"
}
resp = requests.post(
    "http://localhost:8000/v1/audio/transcriptions",
    data=data,
    files=files
)
print(resp.json()["text"])

Key Source Files

Path Purpose
backend/api/routers/openai_compat.py Main router with model mapping and format handling
services/tts_backend.py Backend registration and get_engine_instance_for
services/audio_io.py Low-level audio encoding utilities
services/model_manager.py GPU pooling and admission control
services/text_normalization.py Text preprocessing for synthesis

Summary

  • model="tts-1" or "tts-1-hd" selects the active TTS backend; engine IDs target specific backends directly
  • STT defaults to "whisper-1" mapping to active ASR, with arbitrary engine IDs for precise control
  • Six formats supported: mp3, opus, aac, flac, wav, pcm—with automatic WAV fallback for missing dependencies
  • Singleton engine instances prevent repeated initialization while respecting resource limits through shared admission control

Frequently Asked Questions

What happens if I specify an unknown model name?

The router returns 400 Bad Request with a JSON error body listing all supported engine IDs. This occurs in lines 95–102 of openai_compat.py when get_backend_class fails to find a matching registration.

Why does my mp3 request return a wav file?

The mp3 encoder requires ffmpeg. If ffmpeg is not installed or accessible, the router catches the RuntimeError and falls back to WAV format (lines 221–232). Check your server logs for encoding warnings.

Can I use the same model parameter for both TTS and STT?

No—the namespaces are separate. TTS accepts "tts-1", "tts-1-hd", and TTS engine IDs. STT accepts "whisper-1" and ASR engine IDs. Passing a TTS engine ID to the transcriptions endpoint triggers the unknown model error.

Does the pcm format include a WAV header?

No. The pcm response format returns raw 16-bit little-endian PCM samples without any container header (lines 244–258). You must manually specify sample rate (24kHz for most VoiceStudio engines) and channel count when playing or processing these bytes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →