How to Use the Text-to-Speech (TTS) Endpoint in mlx-omni-server for Audio Generation

The mlx-omni-server exposes an OpenAI-compatible POST /v1/audio/speech endpoint that generates audio from text using either the F5-TTS or mlx-audio backend, streaming the result in formats like WAV, MP3, or FLAC.

The madroidmaq/mlx-omni-server project provides a FastAPI-based server that brings OpenAI-compatible APIs to Apple's MLX framework. One of its core features is the text-to-speech (TTS) endpoint, which converts text input into spoken audio files using local MLX-optimized models.

TTS Endpoint Architecture

The TTS implementation follows a layered architecture that separates routing, validation, and model execution. Understanding this flow helps troubleshoot issues and optimize performance.

Request Flow

When you send a POST request to /v1/audio/speech, the system processes it through four distinct stages:

  1. Router – src/mlx_omni_server/tts/tts.py receives the request and initializes the service
  2. Validation – src/mlx_omni_server/tts/schema.py validates the payload against the TTSRequest Pydantic model
  3. Service Layer – src/mlx_omni_server/tts/tts_service.py selects the appropriate model adapter (F5Model or MlxAudioModel)
  4. Response Streaming – Generated audio bytes are wrapped in a StreamingResponse with the correct MIME type and Content-Disposition headers

Core Components

Component File Path Responsibility
Router src/mlx_omni_server/tts/tts.py Defines the /v1/audio/speech route and handles HTTP responses
Schema src/mlx_omni_server/tts/schema.py Validates model, input, voice, response_format, and speed parameters
Service src/mlx_omni_server/tts/tts_service.py Orchestrates model selection and audio generation
Adapters Same as above F5Model uses f5_tts_mlx.generate; MlxAudioModel uses mlx_audio.tts.generate
API Integration src/mlx_omni_server/routers.py Mounts the TTS router into the global API

Sending Requests to the TTS Endpoint

The server listens on port 10240 by default. You can interact with the endpoint using any HTTP client.

Using cURL

The official documentation in docs/apis/audio.md provides this example:

curl -X POST "http://localhost:10240/v1/audio/speech" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "lucasnewman/f5-tts-mlx",
        "input": "MLX project is awesome",
        "voice": "alloy",
        "response_format": "wav",
        "speed": 1.0
      }' \
  --output speech.wav

This command streams the audio directly to speech.wav. The endpoint supports multiple formats including mp3, opus, aac, flac, wav, and pcm.

Python with Requests

For synchronous Python applications:

import requests

url = "http://localhost:10240/v1/audio/speech"
payload = {
    "model": "lucasnewman/f5-tts-mlx",
    "input": "MLX project is awesome",
    "voice": "alloy",
    "response_format": "wav",
    "speed": 1.0
}

response = requests.post(url, json=payload, stream=True)
response.raise_for_status()

with open("speech.wav", "wb") as f:
    for chunk in response.iter_content(chunk_size=8192):
        f.write(chunk)

Async Python with httpx

For asynchronous workflows:

import httpx

async def generate_speech():
    async with httpx.AsyncClient() as client:
        resp = await client.post(
            "http://localhost:10240/v1/audio/speech",
            json={
                "model": "lucasnewman/f5-tts-mlx",
                "input": "Hello from the TTS endpoint!",
                "voice": "alloy",
                "response_format": "mp3",
                "speed": 1.2,
            },
            timeout=60.0,
        )
        resp.raise_for_status()
        with open("speech.mp3", "wb") as fp:
            fp.write(resp.content)

Request Parameters and Schema

The TTSRequest model in src/mlx_omni_server/tts/schema.py defines the following fields:

  • model – Model identifier string. Use lucasnewman/f5-tts-mlx to trigger the F5-TTS backend; any other value defaults to the generic mlx-audio implementation.
  • input – The text string to convert to speech.
  • voice – Voice identifier. For F5 models, use "alloy"; for mlx-audio, the default is typically "af_sky".
  • response_format – Audio encoding format. Valid options are mp3, opus, aac, flac, wav, and pcm.
  • speed – Playback speed multiplier ranging from 0.25 to 4.0.

Passing Extra Parameters

The schema accepts additional fields beyond the standard OpenAI specification. Any extra parameters included in the JSON payload are captured by request.get_extra_params() and passed directly to the underlying generation library, allowing you to tune model-specific settings.

Model Adapters and Backend Selection

The TTSService class in src/mlx_omni_server/tts/tts_service.py automatically selects the appropriate adapter based on the model parameter:

  • F5Model – Activated when the model string contains f5-tts-mlx. Wraps the f5_tts_mlx.generate function.
  • MlxAudioModel – Used for all other model identifiers. Wraps mlx_audio.tts.generate.

Both adapters implement a generate_audio() method that writes to a temporary file (default sample.wav), which the service reads and streams back to the client before deletion. Ensure the server process has write permissions in the working directory.

Summary

  • The TTS endpoint at POST /v1/audio/speech provides OpenAI-compatible speech synthesis using MLX-optimized models.
  • Two backends are available: F5-TTS (lucasnewman/f5-tts-mlx) and generic mlx-audio, selected automatically based on the model parameter.
  • Six audio formats are supported: mp3, opus, aac, flac, wav, and pcm.
  • The service generates a temporary file on disk before streaming, requiring write permissions in the server working directory.
  • Extra parameters in the request body are forwarded to the underlying TTS library for advanced configuration.

Frequently Asked Questions

What audio formats does the mlx-omni-server TTS endpoint support?

The endpoint supports six formats defined in the AudioFormat enum within src/mlx_omni_server/tts/schema.py: mp3, opus, aac, flac, wav, and pcm. The Content-Type header of the response matches your selected response_format.

How does the server choose between F5-TTS and mlx-audio?

The TTSService class inspects the model parameter in src/mlx_omni_server/tts/tts_service.py. If the string matches lucasnewman/f5-tts-mlx, it instantiates F5Model; otherwise, it uses MlxAudioModel. Each adapter calls its respective underlying library (f5_tts_mlx or mlx_audio).

Can I adjust voice characteristics beyond the speed parameter?

Yes. While speed is the standard parameter, you can include additional fields in your JSON payload. The TTSRequest model captures these via get_extra_params() and passes them directly to the underlying generation function, enabling model-specific adjustments like pitch or speaker embeddings.

Why does the server require write permissions for TTS generation?

Both model adapters write the generated audio to a temporary file named sample.wav (or similar) on the local filesystem before streaming it to the client. The service deletes this file after reading, but the process must have write access to the working directory to create it initially.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →