# How the OpenAI-Compatible Audio API Maps `model` to Engine Selection and Supports Output Formats

> Discover how the VoiceStudio OpenAI-compatible audio API maps model parameters to TTS/STT engine selection and supports six audio output formats with fallback.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: api-reference
- Published: 2026-09-06

---

**The OpenAI-compatible audio router in VoiceStudio translates the `model` parameter into specific TTS or STT engine instances and supports six audio output formats with graceful fallback handling.**

VoiceStudio provides a drop-in replacement for OpenAI's speech endpoints through its `/v1/audio/*` routes. This implementation, found in [`backend/api/routers/openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/openai_compat.py), reuses the platform's native engine architecture while maintaining full compatibility with OpenAI's API contract. Understanding how the `model` parameter drives engine selection and which formats are supported helps developers integrate VoiceStudio seamlessly into existing workflows.

## Model-to-Engine Mapping for Text-to-Speech

When clients POST to `/v1/audio/speech`, the `model` field undergoes resolution through the `_resolve_engine` helper function.

### OpenAI Aliases vs. Direct Engine IDs

The router accepts two categories of `model` values:

| Model value | Mapping behavior |
|-------------|----------------|
| `"tts-1"` or `"tts-1-hd"` | Returns the **active TTS backend**—whatever engine is currently selected by the user in the VoiceStudio UI |
| Valid engine IDs (`"omnivoice"`, `"voxcpm2"`, `"cosyvoice"`, etc.) | Looks up the backend class via `services.tts_backend.get_backend_class` and returns a **singleton instance** through `get_engine_instance_for` |
| Unknown or invalid value | Raises **400 Bad Request** with a list of supported engine IDs |

The resolution logic spans lines 61–102 of [`openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/openai_compat.py). For registered backends, the singleton pattern prevents repeated initialization overhead while still permitting hot-swapping of the active engine.

## Model-to-Engine Mapping for Speech-to-Text

The STT endpoint `/v1/audio/transcriptions` applies parallel logic with Whisper-specific defaults.

The `model` form field defaults to `"whisper-1"`, which maps to the **active ASR backend**. Any other string is interpreted as a specific VoiceStudio ASR engine identifier such as `whisperx` or `faster-whisper` (lines 15–20).

This design ensures that existing OpenAI SDK clients work without modification, while power users can target specific engines by passing their VoiceStudio IDs.

## Supported Output Formats in TTS

The `response_format` field in `SpeechRequest` accepts a **Literal** type with six possible values:

```

"mp3", "opus", "aac", "flac", "wav", "pcm"

```

After synthesis produces a raw `torch.Tensor`, the `_encode_audio` function routes to format-specific encoders:

| Format | Encoding implementation |
|--------|------------------------|
| `wav` | `torchaudio.save(..., format="wav")` → MIME `audio/wav` |
| `flac` | `torchaudio.save(..., format="flac")` → MIME `audio/flac` |
| `mp3` | Attempts `format="mp3"`; falls back to WAV if ffmpeg is unavailable |
| `opus` | Attempts `format="ogg"` (Opus container); falls back to WAV on failure |
| `pcm` | Returns raw 16-bit little-endian PCM bytes without container |
| `aac` | Not directly supported; automatically falls back to WAV |

The encoder selection spans lines 205–262. Error handling is defensive: when preferred containers cannot be written, the router degrades to WAV rather than failing the request. The final `StreamingResponse` carries correct MIME types and file extensions to ensure proper client-side handling.

## Architecture Benefits

Three design decisions make this implementation robust:

- **Engine-agnostic routing** — The same admission-control logic from native endpoints enforces GPU/CPU caps and model-load budgets
- **Singleton backend instances** — Cached per-engine classes eliminate repeated sidecar spawns
- **Graceful degradation** — Missing ffmpeg or unsupported codecs trigger automatic fallback to universally compatible WAV

## Code Examples

### TTS with OpenAI Alias

```python
import requests

resp = requests.post(
    "http://localhost:8000/v1/audio/speech",
    json={
        "model": "tts-1",          # maps to currently active backend

        "input": "Hello, world!",
        "voice": "default",
        "response_format": "mp3"
    },
)
with open("hello.mp3", "wb") as f:
    f.write(resp.content)

```

### STT with Specific Engine

```python
import requests

files = {"file": open("speech.wav", "rb")}
data = {
    "model": "faster-whisper",   # selects faster-whisper backend directly

    "response_format": "text"
}
resp = requests.post(
    "http://localhost:8000/v1/audio/transcriptions",
    data=data,
    files=files
)
print(resp.json()["text"])

```

## Key Source Files

| Path | Purpose |
|------|---------|
| [`backend/api/routers/openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/openai_compat.py) | Main router with model mapping and format handling |
| [`services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/tts_backend.py) | Backend registration and `get_engine_instance_for` |
| [`services/audio_io.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/audio_io.py) | Low-level audio encoding utilities |
| [`services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/model_manager.py) | GPU pooling and admission control |
| [`services/text_normalization.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/text_normalization.py) | Text preprocessing for synthesis |

## Summary

- **`model="tts-1"`** or **`"tts-1-hd"`** selects the active TTS backend; engine IDs target specific backends directly
- **STT defaults** to `"whisper-1"` mapping to active ASR, with arbitrary engine IDs for precise control
- **Six formats supported**: mp3, opus, aac, flac, wav, pcm—with automatic WAV fallback for missing dependencies
- **Singleton engine instances** prevent repeated initialization while respecting resource limits through shared admission control

## Frequently Asked Questions

### What happens if I specify an unknown model name?

The router returns **400 Bad Request** with a JSON error body listing all supported engine IDs. This occurs in lines 95–102 of [`openai_compat.py`](https://github.com/debpalash/VoiceStudio/blob/main/openai_compat.py) when `get_backend_class` fails to find a matching registration.

### Why does my mp3 request return a wav file?

The mp3 encoder requires ffmpeg. If ffmpeg is not installed or accessible, the router catches the `RuntimeError` and falls back to WAV format (lines 221–232). Check your server logs for encoding warnings.

### Can I use the same model parameter for both TTS and STT?

No—the namespaces are separate. TTS accepts `"tts-1"`, `"tts-1-hd"`, and TTS engine IDs. STT accepts `"whisper-1"` and ASR engine IDs. Passing a TTS engine ID to the transcriptions endpoint triggers the unknown model error.

### Does the pcm format include a WAV header?

No. The `pcm` response format returns **raw 16-bit little-endian PCM samples** without any container header (lines 244–258). You must manually specify sample rate (24kHz for most VoiceStudio engines) and channel count when playing or processing these bytes.