How to Call the OpenAI-Compatible /v1/audio/speech Endpoint with Supertonic Serve

You can call the OpenAI-compatible /v1/audio/speech endpoint by running supertonic serve to start a local FastAPI server, then sending a POST request with a JSON payload containing model, input, voice, and optional speed and lang parameters to receive synthesized audio binary.

The supertone-inc/supertonic repository provides an open-source text-to-speech engine that exposes a standards-compliant HTTP API. When you execute supertonic serve, the application launches a local server implemented in py/helper.py that mirrors the OpenAI /v1/audio/speech specification, enabling you to use existing OpenAI client libraries or direct HTTP requests to generate speech locally without cloud dependencies.

Understanding the Endpoint Architecture

The FastAPI application defined in py/helper.py exposes two TTS routes: /v1/tts (native Supertonic API) and /v1/audio/speech (OpenAI-compatible). The latter route strictly follows the OpenAI audio speech request and response schema, parsing the JSON body to extract synthesis parameters and returning raw audio bytes. This architecture allows any client configured for the OpenAI API to communicate with your local Supertonic instance by simply changing the base URL.

Starting the Supertonic Server

Before invoking the endpoint, you must install the optional serve dependencies and launch the server. According to the repository documentation in README.md and the entry point defined in py/__init__.py, use the following commands:

Install the serve extras:

pip install "supertonic[serve]"

Start the server on your desired host and port:

supertonic serve --host 127.0.0.1 --port 7788

The server will listen at http://127.0.0.1:7788 and provide interactive OpenAPI documentation at /docs.

Calling the /v1/audio/speech Endpoint

Direct HTTP Request with curl

Send a POST request to the local endpoint with the required JSON payload:

curl -X POST "http://127.0.0.1:7788/v1/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "supertonic-tts",
    "input": "Hello from the Supertonic local server",
    "voice": "M1",
    "speed": 1.0,
    "lang": "en"
  }' \
  --output speech.wav

The --output flag saves the binary audio response (WAV or MP3) to disk.

Using the Official OpenAI Python Client

Because the endpoint is fully compatible with OpenAI's specification, you can redirect the official Python client to your local Supertonic server:

import openai

openai.api_base = "http://127.0.0.1:7788/v1"
openai.api_key = "not-needed"  # any non-empty string satisfies the library

response = openai.Audio.speech.create(
    model="supertonic-tts",
    input="Testing the OpenAI-compatible endpoint",
    voice="F2",
    speed=1.2,
    lang="en"
)

with open("output.wav", "wb") as f:
    f.write(response.content)

Using Python requests

The repository includes py/example_pypi.py, which demonstrates a native implementation using the requests library:

import requests

url = "http://127.0.0.1:7788/v1/audio/speech"
payload = {
    "model": "supertonic-tts",
    "input": "Supertonic runs locally without external APIs",
    "voice": "M1",
    "lang": "en"
}
headers = {"Content-Type": "application/json"}

r = requests.post(url, json=payload, headers=headers)
with open("result.wav", "wb") as f:
    f.write(r.content)

Request Payload Specifications

The OpenAI-compatible endpoint expects a JSON object with the following fields:

  • model: Must be "supertonic-tts" (the only supported model identifier).
  • input: The text string to synthesize into speech.
  • voice: Voice style identifier (e.g., "M1", "F2").
  • speed: Optional playback speed multiplier (range approximately 0.7 to 2.0, default 1.0).
  • lang: Optional language code (e.g., "en", "ja", "ru", or "na" for language-agnostic mode).

Response Handling

The server returns standardized HTTP status codes:

  • 200 OK: Successful synthesis returning binary audio data with Content-Type: audio/mpeg or audio/wav.
  • 400/422: Validation errors returned in a JSON format matching OpenAI's error schema, typically indicating missing required fields or invalid parameter values.

Summary

  • supertonic serve starts a local FastAPI server defined in py/helper.py that exposes the OpenAI-compatible /v1/audio/speech endpoint.
  • The endpoint accepts standard OpenAI request payloads containing model, input, voice, and optional speed and lang parameters.
  • You can invoke the endpoint using curl, the official OpenAI Python client, or direct HTTP libraries like requests as shown in py/example_pypi.py.
  • The server returns binary audio data (WAV or MP3) on success, making it compatible with existing audio playback pipelines.

Frequently Asked Questions

What authentication does the Supertonic OpenAI-compatible endpoint require?

The local server does not implement authentication. When using the OpenAI Python client, you must provide any non-empty string for api_key (e.g., "not-needed") to satisfy the library's requirements, as the server ignores the Authorization header.

Can I use voice settings beyond speed and language?

The OpenAI-compatible endpoint strictly supports the standard OpenAI parameters: voice, speed, and lang. Additional Supertonic-specific features are available through the native /v1/tts endpoint, but these are not accessible via the OpenAI-compatible route.

Why am I receiving a 422 validation error?

HTTP 422 errors indicate that your JSON payload failed Pydantic validation. Ensure you include the required fields (model, input, voice), that model is exactly "supertonic-tts", and that numeric values like speed fall within the supported range (approximately 0.7 to 2.0).

How do I change the audio output format?

The response format (WAV or MP3) is determined by the Supertonic engine configuration and model settings. The endpoint automatically sets the Content-Type header to audio/wav or audio/mpeg based on the generated output; you cannot specify the format via the OpenAI-compatible request parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →