How to Use the Speech-to-Text (Whisper) Endpoint in mlx-omni-server

The mlx-omni-server exposes Whisper speech-to-text capabilities via POST /audio/transcriptions and POST /v1/audio/transcriptions, accepting audio files and returning transcriptions in multiple formats including JSON, SRT, VTT, and plain text.

The mlx-omni-server project provides OpenAI-compatible API endpoints for MLX-based models. Its speech-to-text implementation leverages Apple's MLX framework to run Whisper inference locally, offering a privacy-focused alternative to cloud-based transcription services while maintaining API compatibility with OpenAI's audio transcriptions specification.

Endpoint Routes and Architecture

The STT service is implemented in src/mlx_omni_server/stt/stt.py and mounted via src/mlx_omni_server/routers.py. Two identical endpoints handle incoming requests:

  • POST /audio/transcriptions
  • POST /v1/audio/transcriptions

Both routes utilize the STTService class defined in src/mlx_omni_server/stt/whisper_model.py, which orchestrates temporary file handling and invokes the underlying Whisper model.

Request Validation and Parameters

Incoming requests are validated through the STTRequestForm class located in src/mlx_omni_server/stt/schema.py. This form performs several critical validations before processing begins:

  • File type verification for uploaded audio files
  • Temperature range constraints (typically 0.0 to 1.0)
  • Language code validation (e.g., en, es, fr)
  • Compatibility checks between timestamp_granularities and response_format

The form normalizes the timestamp_granularities[] form field into a list of TimestampGranularity enum values. Valid options include segment for segment-level timestamps and word for word-level timestamps.

Supported Response Formats

The endpoint supports five distinct output formats controlled by the response_format parameter:

Format FastAPI Response Type Content Description
json JSONResponse Simple JSON object with {"text": "transcription"}
verbose_json JSONResponse Complete Whisper output including segments, words, duration, and metadata
text PlainTextResponse Raw transcription text only
srt Response with headers SubRip subtitle format with Content-Disposition for download
vtt Response with headers WebVTT subtitle format with Content-Disposition for download

Important: Word-level timestamps (timestamp_granularities[]=word) require response_format=verbose_json. The STTRequestForm enforces this compatibility constraint during validation.

Implementation Workflow

When processing a request, the STTService executes the following workflow:

  1. Saves the uploaded audio file to a temporary location on disk
  2. Calls mlx_whisper.transcribe with the specified model name, language, temperature, optional prompt, and timestamp granularity flags
  3. Formats the raw inference results according to the requested response_format
  4. Returns the appropriate FastAPI response type based on the format selection

Practical Code Examples

cURL Request

curl -X POST "http://localhost:8000/audio/transcriptions" \
  -H "Accept: application/json" \
  -F "file=@/path/to/audio.wav" \
  -F "model=mlx/whisper-base" \
  -F "language=en" \
  -F "response_format=json" \
  -F "temperature=0.0" \
  -F "timestamp_granularities[]=segment"

Python with requests

import requests

url = "http://localhost:8000/v1/audio/transcriptions"
files = {"file": open("speech.mp3", "rb")}
data = {
    "model": "mlx/whisper-base",
    "language": "en",
    "response_format": "verbose_json",
    "temperature": "0.0",
    "timestamp_granularities[]": "word",  # Requires verbose_json

}

resp = requests.post(url, files=files, data=data)
result = resp.json()  # Contains full segments and word-level timestamps

Python with httpx (Async)

import httpx

async def transcribe():
    async with httpx.AsyncClient(base_url="http://localhost:8000") as client:
        files = {"file": ("audio.wav", open("audio.wav", "rb"), "audio/wav")}
        data = {
            "model": "mlx/whisper-large",
            "response_format": "srt",
            "timestamp_granularities[]": "segment",
        }
        r = await client.post("/audio/transcriptions", files=files, data=data)
        r.raise_for_status()
        print(r.text)  # SRT subtitle content

# asyncio.run(transcribe())

Verbose JSON Response Structure

When requesting verbose_json, the endpoint returns a comprehensive payload including timing metadata:

{
  "task": "transcribe",
  "language": "en",
  "duration": 12.34,
  "text": "Hello world, this is a test.",
  "words": [
    {"word": "Hello", "start": 0.0, "end": 0.5},
    {"word": "world", "start": 0.5, "end": 0.9}
  ],
  "segments": [
    {
      "id": 0,
      "seek": 0,
      "start": 0.0,
      "end": 2.1,
      "text": "Hello world",
      "tokens": [50364, 292, 1234],
      "temperature": 0.0,
      "avg_logprob": -0.02,
      "compression_ratio": 1.2,
      "no_speech_prob": 0.01
    }
  ]
}

Summary

  • The mlx-omni-server provides OpenAI-compatible speech-to-text endpoints at /audio/transcriptions and /v1/audio/transcriptions
  • Request validation occurs through STTRequestForm in src/mlx_omni_server/stt/schema.py, ensuring parameter compatibility
  • The service supports five response formats: json, verbose_json, text, srt, and vtt
  • Word-level timestamps require verbose_json format with timestamp_granularities[]=word
  • Underlying inference uses mlx_whisper.transcribe as implemented in src/mlx_omni_server/stt/whisper_model.py

Frequently Asked Questions

What audio file formats does the Whisper endpoint accept?

The endpoint accepts standard audio formats supported by the underlying MLX Whisper implementation. The STTRequestForm in src/mlx_omni_server/stt/schema.py validates file types upon upload, ensuring compatibility before the audio reaches the mlx_whisper.transcribe call in src/mlx_omni_server/stt/whisper_model.py.

How do I enable word-level timestamps in the transcription?

Set response_format to verbose_json and include timestamp_granularities[]=word in your request parameters. The validation logic in STTRequestForm enforces that word-level granularities are only requested with verbose JSON format, as this is the only format that includes the detailed words array in the response structure.

Can I use this endpoint as a drop-in replacement for OpenAI's Whisper API?

Yes, the endpoint routes and request parameters mirror the OpenAI Audio Transcriptions API specification. However, you must use MLX-specific model names with the mlx/ prefix (e.g., mlx/whisper-base or mlx/whisper-large) as defined in the STTService implementation in src/mlx_omni_server/stt/whisper_model.py.

Where is the transcription logic implemented in the source code?

The core transcription logic resides in src/mlx_omni_server/stt/whisper_model.py, specifically within the STTService class that handles file storage and calls mlx_whisper.transcribe. API routing and response formatting are defined in src/mlx_omni_server/stt/stt.py, with the router registered in the main application via src/mlx_omni_server/routers.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →