How the OpenAI-Compatible Audio API Maps `model` to Engine Selection and Supports Output Formats
The OpenAI-compatible audio router in VoiceStudio translates the model parameter into specific TTS or STT engine instances and supports six audio output formats with graceful fallback handling.
VoiceStudio provides a drop-in replacement for OpenAI's speech endpoints through its /v1/audio/* routes. This implementation, found in backend/api/routers/openai_compat.py, reuses the platform's native engine architecture while maintaining full compatibility with OpenAI's API contract. Understanding how the model parameter drives engine selection and which formats are supported helps developers integrate VoiceStudio seamlessly into existing workflows.
Model-to-Engine Mapping for Text-to-Speech
When clients POST to /v1/audio/speech, the model field undergoes resolution through the _resolve_engine helper function.
OpenAI Aliases vs. Direct Engine IDs
The router accepts two categories of model values:
| Model value | Mapping behavior |
|---|---|
"tts-1" or "tts-1-hd" |
Returns the active TTS backend—whatever engine is currently selected by the user in the VoiceStudio UI |
Valid engine IDs ("omnivoice", "voxcpm2", "cosyvoice", etc.) |
Looks up the backend class via services.tts_backend.get_backend_class and returns a singleton instance through get_engine_instance_for |
| Unknown or invalid value | Raises 400 Bad Request with a list of supported engine IDs |
The resolution logic spans lines 61–102 of openai_compat.py. For registered backends, the singleton pattern prevents repeated initialization overhead while still permitting hot-swapping of the active engine.
Model-to-Engine Mapping for Speech-to-Text
The STT endpoint /v1/audio/transcriptions applies parallel logic with Whisper-specific defaults.
The model form field defaults to "whisper-1", which maps to the active ASR backend. Any other string is interpreted as a specific VoiceStudio ASR engine identifier such as whisperx or faster-whisper (lines 15–20).
This design ensures that existing OpenAI SDK clients work without modification, while power users can target specific engines by passing their VoiceStudio IDs.
Supported Output Formats in TTS
The response_format field in SpeechRequest accepts a Literal type with six possible values:
"mp3", "opus", "aac", "flac", "wav", "pcm"
After synthesis produces a raw torch.Tensor, the _encode_audio function routes to format-specific encoders:
| Format | Encoding implementation |
|---|---|
wav |
torchaudio.save(..., format="wav") → MIME audio/wav |
flac |
torchaudio.save(..., format="flac") → MIME audio/flac |
mp3 |
Attempts format="mp3"; falls back to WAV if ffmpeg is unavailable |
opus |
Attempts format="ogg" (Opus container); falls back to WAV on failure |
pcm |
Returns raw 16-bit little-endian PCM bytes without container |
aac |
Not directly supported; automatically falls back to WAV |
The encoder selection spans lines 205–262. Error handling is defensive: when preferred containers cannot be written, the router degrades to WAV rather than failing the request. The final StreamingResponse carries correct MIME types and file extensions to ensure proper client-side handling.
Architecture Benefits
Three design decisions make this implementation robust:
- Engine-agnostic routing — The same admission-control logic from native endpoints enforces GPU/CPU caps and model-load budgets
- Singleton backend instances — Cached per-engine classes eliminate repeated sidecar spawns
- Graceful degradation — Missing ffmpeg or unsupported codecs trigger automatic fallback to universally compatible WAV
Code Examples
TTS with OpenAI Alias
import requests
resp = requests.post(
"http://localhost:8000/v1/audio/speech",
json={
"model": "tts-1", # maps to currently active backend
"input": "Hello, world!",
"voice": "default",
"response_format": "mp3"
},
)
with open("hello.mp3", "wb") as f:
f.write(resp.content)
STT with Specific Engine
import requests
files = {"file": open("speech.wav", "rb")}
data = {
"model": "faster-whisper", # selects faster-whisper backend directly
"response_format": "text"
}
resp = requests.post(
"http://localhost:8000/v1/audio/transcriptions",
data=data,
files=files
)
print(resp.json()["text"])
Key Source Files
| Path | Purpose |
|---|---|
backend/api/routers/openai_compat.py |
Main router with model mapping and format handling |
services/tts_backend.py |
Backend registration and get_engine_instance_for |
services/audio_io.py |
Low-level audio encoding utilities |
services/model_manager.py |
GPU pooling and admission control |
services/text_normalization.py |
Text preprocessing for synthesis |
Summary
model="tts-1"or"tts-1-hd"selects the active TTS backend; engine IDs target specific backends directly- STT defaults to
"whisper-1"mapping to active ASR, with arbitrary engine IDs for precise control - Six formats supported: mp3, opus, aac, flac, wav, pcm—with automatic WAV fallback for missing dependencies
- Singleton engine instances prevent repeated initialization while respecting resource limits through shared admission control
Frequently Asked Questions
What happens if I specify an unknown model name?
The router returns 400 Bad Request with a JSON error body listing all supported engine IDs. This occurs in lines 95–102 of openai_compat.py when get_backend_class fails to find a matching registration.
Why does my mp3 request return a wav file?
The mp3 encoder requires ffmpeg. If ffmpeg is not installed or accessible, the router catches the RuntimeError and falls back to WAV format (lines 221–232). Check your server logs for encoding warnings.
Can I use the same model parameter for both TTS and STT?
No—the namespaces are separate. TTS accepts "tts-1", "tts-1-hd", and TTS engine IDs. STT accepts "whisper-1" and ASR engine IDs. Passing a TTS engine ID to the transcriptions endpoint triggers the unknown model error.
Does the pcm format include a WAV header?
No. The pcm response format returns raw 16-bit little-endian PCM samples without any container header (lines 244–258). You must manually specify sample rate (24kHz for most VoiceStudio engines) and channel count when playing or processing these bytes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →