# How to Call the OpenAI-Compatible /v1/audio/speech Endpoint with Supertonic Serve

> Easily call the OpenAI compatible audio speech endpoint using supertonic serve. Send a POST request to your local FastAPI server and get synthesized audio binary.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-14

---

**You can call the OpenAI-compatible `/v1/audio/speech` endpoint by running `supertonic serve` to start a local FastAPI server, then sending a POST request with a JSON payload containing `model`, `input`, `voice`, and optional `speed` and `lang` parameters to receive synthesized audio binary.**

The `supertone-inc/supertonic` repository provides an open-source text-to-speech engine that exposes a standards-compliant HTTP API. When you execute `supertonic serve`, the application launches a local server implemented in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) that mirrors the OpenAI `/v1/audio/speech` specification, enabling you to use existing OpenAI client libraries or direct HTTP requests to generate speech locally without cloud dependencies.

## Understanding the Endpoint Architecture

The FastAPI application defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) exposes two TTS routes: `/v1/tts` (native Supertonic API) and `/v1/audio/speech` (OpenAI-compatible). The latter route strictly follows the OpenAI audio speech request and response schema, parsing the JSON body to extract synthesis parameters and returning raw audio bytes. This architecture allows any client configured for the OpenAI API to communicate with your local Supertonic instance by simply changing the base URL.

## Starting the Supertonic Server

Before invoking the endpoint, you must install the optional serve dependencies and launch the server. According to the repository documentation in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) and the entry point defined in [`py/__init__.py`](https://github.com/supertone-inc/supertonic/blob/main/py/__init__.py), use the following commands:

Install the serve extras:

```bash
pip install "supertonic[serve]"

```

Start the server on your desired host and port:

```bash
supertonic serve --host 127.0.0.1 --port 7788

```

The server will listen at `http://127.0.0.1:7788` and provide interactive OpenAPI documentation at `/docs`.

## Calling the /v1/audio/speech Endpoint

### Direct HTTP Request with curl

Send a POST request to the local endpoint with the required JSON payload:

```bash
curl -X POST "http://127.0.0.1:7788/v1/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "supertonic-tts",
    "input": "Hello from the Supertonic local server",
    "voice": "M1",
    "speed": 1.0,
    "lang": "en"
  }' \
  --output speech.wav

```

The `--output` flag saves the binary audio response (WAV or MP3) to disk.

### Using the Official OpenAI Python Client

Because the endpoint is fully compatible with OpenAI's specification, you can redirect the official Python client to your local Supertonic server:

```python
import openai

openai.api_base = "http://127.0.0.1:7788/v1"
openai.api_key = "not-needed"  # any non-empty string satisfies the library

response = openai.Audio.speech.create(
    model="supertonic-tts",
    input="Testing the OpenAI-compatible endpoint",
    voice="F2",
    speed=1.2,
    lang="en"
)

with open("output.wav", "wb") as f:
    f.write(response.content)

```

### Using Python requests

The repository includes [`py/example_pypi.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_pypi.py), which demonstrates a native implementation using the `requests` library:

```python
import requests

url = "http://127.0.0.1:7788/v1/audio/speech"
payload = {
    "model": "supertonic-tts",
    "input": "Supertonic runs locally without external APIs",
    "voice": "M1",
    "lang": "en"
}
headers = {"Content-Type": "application/json"}

r = requests.post(url, json=payload, headers=headers)
with open("result.wav", "wb") as f:
    f.write(r.content)

```

## Request Payload Specifications

The OpenAI-compatible endpoint expects a JSON object with the following fields:

- **model**: Must be `"supertonic-tts"` (the only supported model identifier).
- **input**: The text string to synthesize into speech.
- **voice**: Voice style identifier (e.g., `"M1"`, `"F2"`).
- **speed**: Optional playback speed multiplier (range approximately **0.7 to 2.0**, default **1.0**).
- **lang**: Optional language code (e.g., `"en"`, `"ja"`, `"ru"`, or `"na"` for language-agnostic mode).

## Response Handling

The server returns standardized HTTP status codes:

- **200 OK**: Successful synthesis returning binary audio data with `Content-Type: audio/mpeg` or `audio/wav`.
- **400/422**: Validation errors returned in a JSON format matching OpenAI's error schema, typically indicating missing required fields or invalid parameter values.

## Summary

- **`supertonic serve`** starts a local FastAPI server defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) that exposes the OpenAI-compatible **`/v1/audio/speech`** endpoint.
- The endpoint accepts standard OpenAI request payloads containing `model`, `input`, `voice`, and optional `speed` and `lang` parameters.
- You can invoke the endpoint using **curl**, the **official OpenAI Python client**, or direct HTTP libraries like **requests** as shown in [`py/example_pypi.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_pypi.py).
- The server returns binary audio data (WAV or MP3) on success, making it compatible with existing audio playback pipelines.

## Frequently Asked Questions

### What authentication does the Supertonic OpenAI-compatible endpoint require?

The local server does not implement authentication. When using the OpenAI Python client, you must provide any non-empty string for `api_key` (e.g., `"not-needed"`) to satisfy the library's requirements, as the server ignores the Authorization header.

### Can I use voice settings beyond speed and language?

The OpenAI-compatible endpoint strictly supports the standard OpenAI parameters: `voice`, `speed`, and `lang`. Additional Supertonic-specific features are available through the native `/v1/tts` endpoint, but these are not accessible via the OpenAI-compatible route.

### Why am I receiving a 422 validation error?

HTTP 422 errors indicate that your JSON payload failed Pydantic validation. Ensure you include the required fields (`model`, `input`, `voice`), that `model` is exactly `"supertonic-tts"`, and that numeric values like `speed` fall within the supported range (approximately 0.7 to 2.0).

### How do I change the audio output format?

The response format (WAV or MP3) is determined by the Supertonic engine configuration and model settings. The endpoint automatically sets the `Content-Type` header to `audio/wav` or `audio/mpeg` based on the generated output; you cannot specify the format via the OpenAI-compatible request parameters.