How to Use the Text-to-Speech (TTS) Endpoint in mlx-omni-server for Audio Generation
The mlx-omni-server exposes an OpenAI-compatible POST /v1/audio/speech endpoint that generates audio from text using either the F5-TTS or mlx-audio backend, streaming the result in formats like WAV, MP3, or FLAC.
The madroidmaq/mlx-omni-server project provides a FastAPI-based server that brings OpenAI-compatible APIs to Apple's MLX framework. One of its core features is the text-to-speech (TTS) endpoint, which converts text input into spoken audio files using local MLX-optimized models.
TTS Endpoint Architecture
The TTS implementation follows a layered architecture that separates routing, validation, and model execution. Understanding this flow helps troubleshoot issues and optimize performance.
Request Flow
When you send a POST request to /v1/audio/speech, the system processes it through four distinct stages:
- Router –
src/mlx_omni_server/tts/tts.pyreceives the request and initializes the service - Validation –
src/mlx_omni_server/tts/schema.pyvalidates the payload against theTTSRequestPydantic model - Service Layer –
src/mlx_omni_server/tts/tts_service.pyselects the appropriate model adapter (F5ModelorMlxAudioModel) - Response Streaming – Generated audio bytes are wrapped in a
StreamingResponsewith the correct MIME type andContent-Dispositionheaders
Core Components
| Component | File Path | Responsibility |
|---|---|---|
| Router | src/mlx_omni_server/tts/tts.py |
Defines the /v1/audio/speech route and handles HTTP responses |
| Schema | src/mlx_omni_server/tts/schema.py |
Validates model, input, voice, response_format, and speed parameters |
| Service | src/mlx_omni_server/tts/tts_service.py |
Orchestrates model selection and audio generation |
| Adapters | Same as above | F5Model uses f5_tts_mlx.generate; MlxAudioModel uses mlx_audio.tts.generate |
| API Integration | src/mlx_omni_server/routers.py |
Mounts the TTS router into the global API |
Sending Requests to the TTS Endpoint
The server listens on port 10240 by default. You can interact with the endpoint using any HTTP client.
Using cURL
The official documentation in docs/apis/audio.md provides this example:
curl -X POST "http://localhost:10240/v1/audio/speech" \
-H "Content-Type: application/json" \
-d '{
"model": "lucasnewman/f5-tts-mlx",
"input": "MLX project is awesome",
"voice": "alloy",
"response_format": "wav",
"speed": 1.0
}' \
--output speech.wav
This command streams the audio directly to speech.wav. The endpoint supports multiple formats including mp3, opus, aac, flac, wav, and pcm.
Python with Requests
For synchronous Python applications:
import requests
url = "http://localhost:10240/v1/audio/speech"
payload = {
"model": "lucasnewman/f5-tts-mlx",
"input": "MLX project is awesome",
"voice": "alloy",
"response_format": "wav",
"speed": 1.0
}
response = requests.post(url, json=payload, stream=True)
response.raise_for_status()
with open("speech.wav", "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
Async Python with httpx
For asynchronous workflows:
import httpx
async def generate_speech():
async with httpx.AsyncClient() as client:
resp = await client.post(
"http://localhost:10240/v1/audio/speech",
json={
"model": "lucasnewman/f5-tts-mlx",
"input": "Hello from the TTS endpoint!",
"voice": "alloy",
"response_format": "mp3",
"speed": 1.2,
},
timeout=60.0,
)
resp.raise_for_status()
with open("speech.mp3", "wb") as fp:
fp.write(resp.content)
Request Parameters and Schema
The TTSRequest model in src/mlx_omni_server/tts/schema.py defines the following fields:
model– Model identifier string. Uselucasnewman/f5-tts-mlxto trigger the F5-TTS backend; any other value defaults to the generic mlx-audio implementation.input– The text string to convert to speech.voice– Voice identifier. For F5 models, use"alloy"; for mlx-audio, the default is typically"af_sky".response_format– Audio encoding format. Valid options aremp3,opus,aac,flac,wav, andpcm.speed– Playback speed multiplier ranging from0.25to4.0.
Passing Extra Parameters
The schema accepts additional fields beyond the standard OpenAI specification. Any extra parameters included in the JSON payload are captured by request.get_extra_params() and passed directly to the underlying generation library, allowing you to tune model-specific settings.
Model Adapters and Backend Selection
The TTSService class in src/mlx_omni_server/tts/tts_service.py automatically selects the appropriate adapter based on the model parameter:
F5Model– Activated when the model string containsf5-tts-mlx. Wraps thef5_tts_mlx.generatefunction.MlxAudioModel– Used for all other model identifiers. Wrapsmlx_audio.tts.generate.
Both adapters implement a generate_audio() method that writes to a temporary file (default sample.wav), which the service reads and streams back to the client before deletion. Ensure the server process has write permissions in the working directory.
Summary
- The TTS endpoint at
POST /v1/audio/speechprovides OpenAI-compatible speech synthesis using MLX-optimized models. - Two backends are available: F5-TTS (
lucasnewman/f5-tts-mlx) and generic mlx-audio, selected automatically based on themodelparameter. - Six audio formats are supported:
mp3,opus,aac,flac,wav, andpcm. - The service generates a temporary file on disk before streaming, requiring write permissions in the server working directory.
- Extra parameters in the request body are forwarded to the underlying TTS library for advanced configuration.
Frequently Asked Questions
What audio formats does the mlx-omni-server TTS endpoint support?
The endpoint supports six formats defined in the AudioFormat enum within src/mlx_omni_server/tts/schema.py: mp3, opus, aac, flac, wav, and pcm. The Content-Type header of the response matches your selected response_format.
How does the server choose between F5-TTS and mlx-audio?
The TTSService class inspects the model parameter in src/mlx_omni_server/tts/tts_service.py. If the string matches lucasnewman/f5-tts-mlx, it instantiates F5Model; otherwise, it uses MlxAudioModel. Each adapter calls its respective underlying library (f5_tts_mlx or mlx_audio).
Can I adjust voice characteristics beyond the speed parameter?
Yes. While speed is the standard parameter, you can include additional fields in your JSON payload. The TTSRequest model captures these via get_extra_params() and passes them directly to the underlying generation function, enabling model-specific adjustments like pitch or speaker embeddings.
Why does the server require write permissions for TTS generation?
Both model adapters write the generated audio to a temporary file named sample.wav (or similar) on the local filesystem before streaming it to the client. The service deletes this file after reading, but the process must have write access to the working directory to create it initially.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →