How to Use the Speech-to-Text (Whisper) Endpoint in mlx-omni-server
The mlx-omni-server exposes Whisper speech-to-text capabilities via POST /audio/transcriptions and POST /v1/audio/transcriptions, accepting audio files and returning transcriptions in multiple formats including JSON, SRT, VTT, and plain text.
The mlx-omni-server project provides OpenAI-compatible API endpoints for MLX-based models. Its speech-to-text implementation leverages Apple's MLX framework to run Whisper inference locally, offering a privacy-focused alternative to cloud-based transcription services while maintaining API compatibility with OpenAI's audio transcriptions specification.
Endpoint Routes and Architecture
The STT service is implemented in src/mlx_omni_server/stt/stt.py and mounted via src/mlx_omni_server/routers.py. Two identical endpoints handle incoming requests:
POST /audio/transcriptionsPOST /v1/audio/transcriptions
Both routes utilize the STTService class defined in src/mlx_omni_server/stt/whisper_model.py, which orchestrates temporary file handling and invokes the underlying Whisper model.
Request Validation and Parameters
Incoming requests are validated through the STTRequestForm class located in src/mlx_omni_server/stt/schema.py. This form performs several critical validations before processing begins:
- File type verification for uploaded audio files
- Temperature range constraints (typically 0.0 to 1.0)
- Language code validation (e.g.,
en,es,fr) - Compatibility checks between
timestamp_granularitiesandresponse_format
The form normalizes the timestamp_granularities[] form field into a list of TimestampGranularity enum values. Valid options include segment for segment-level timestamps and word for word-level timestamps.
Supported Response Formats
The endpoint supports five distinct output formats controlled by the response_format parameter:
| Format | FastAPI Response Type | Content Description |
|---|---|---|
json |
JSONResponse |
Simple JSON object with {"text": "transcription"} |
verbose_json |
JSONResponse |
Complete Whisper output including segments, words, duration, and metadata |
text |
PlainTextResponse |
Raw transcription text only |
srt |
Response with headers |
SubRip subtitle format with Content-Disposition for download |
vtt |
Response with headers |
WebVTT subtitle format with Content-Disposition for download |
Important: Word-level timestamps (timestamp_granularities[]=word) require response_format=verbose_json. The STTRequestForm enforces this compatibility constraint during validation.
Implementation Workflow
When processing a request, the STTService executes the following workflow:
- Saves the uploaded audio file to a temporary location on disk
- Calls
mlx_whisper.transcribewith the specified model name, language, temperature, optional prompt, and timestamp granularity flags - Formats the raw inference results according to the requested
response_format - Returns the appropriate FastAPI response type based on the format selection
Practical Code Examples
cURL Request
curl -X POST "http://localhost:8000/audio/transcriptions" \
-H "Accept: application/json" \
-F "file=@/path/to/audio.wav" \
-F "model=mlx/whisper-base" \
-F "language=en" \
-F "response_format=json" \
-F "temperature=0.0" \
-F "timestamp_granularities[]=segment"
Python with requests
import requests
url = "http://localhost:8000/v1/audio/transcriptions"
files = {"file": open("speech.mp3", "rb")}
data = {
"model": "mlx/whisper-base",
"language": "en",
"response_format": "verbose_json",
"temperature": "0.0",
"timestamp_granularities[]": "word", # Requires verbose_json
}
resp = requests.post(url, files=files, data=data)
result = resp.json() # Contains full segments and word-level timestamps
Python with httpx (Async)
import httpx
async def transcribe():
async with httpx.AsyncClient(base_url="http://localhost:8000") as client:
files = {"file": ("audio.wav", open("audio.wav", "rb"), "audio/wav")}
data = {
"model": "mlx/whisper-large",
"response_format": "srt",
"timestamp_granularities[]": "segment",
}
r = await client.post("/audio/transcriptions", files=files, data=data)
r.raise_for_status()
print(r.text) # SRT subtitle content
# asyncio.run(transcribe())
Verbose JSON Response Structure
When requesting verbose_json, the endpoint returns a comprehensive payload including timing metadata:
{
"task": "transcribe",
"language": "en",
"duration": 12.34,
"text": "Hello world, this is a test.",
"words": [
{"word": "Hello", "start": 0.0, "end": 0.5},
{"word": "world", "start": 0.5, "end": 0.9}
],
"segments": [
{
"id": 0,
"seek": 0,
"start": 0.0,
"end": 2.1,
"text": "Hello world",
"tokens": [50364, 292, 1234],
"temperature": 0.0,
"avg_logprob": -0.02,
"compression_ratio": 1.2,
"no_speech_prob": 0.01
}
]
}
Summary
- The mlx-omni-server provides OpenAI-compatible speech-to-text endpoints at
/audio/transcriptionsand/v1/audio/transcriptions - Request validation occurs through
STTRequestForminsrc/mlx_omni_server/stt/schema.py, ensuring parameter compatibility - The service supports five response formats:
json,verbose_json,text,srt, andvtt - Word-level timestamps require
verbose_jsonformat withtimestamp_granularities[]=word - Underlying inference uses
mlx_whisper.transcribeas implemented insrc/mlx_omni_server/stt/whisper_model.py
Frequently Asked Questions
What audio file formats does the Whisper endpoint accept?
The endpoint accepts standard audio formats supported by the underlying MLX Whisper implementation. The STTRequestForm in src/mlx_omni_server/stt/schema.py validates file types upon upload, ensuring compatibility before the audio reaches the mlx_whisper.transcribe call in src/mlx_omni_server/stt/whisper_model.py.
How do I enable word-level timestamps in the transcription?
Set response_format to verbose_json and include timestamp_granularities[]=word in your request parameters. The validation logic in STTRequestForm enforces that word-level granularities are only requested with verbose JSON format, as this is the only format that includes the detailed words array in the response structure.
Can I use this endpoint as a drop-in replacement for OpenAI's Whisper API?
Yes, the endpoint routes and request parameters mirror the OpenAI Audio Transcriptions API specification. However, you must use MLX-specific model names with the mlx/ prefix (e.g., mlx/whisper-base or mlx/whisper-large) as defined in the STTService implementation in src/mlx_omni_server/stt/whisper_model.py.
Where is the transcription logic implemented in the source code?
The core transcription logic resides in src/mlx_omni_server/stt/whisper_model.py, specifically within the STTService class that handles file storage and calls mlx_whisper.transcribe. API routing and response formatting are defined in src/mlx_omni_server/stt/stt.py, with the router registered in the main application via src/mlx_omni_server/routers.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →