# How to Use the Speech-to-Text (Whisper) Endpoint in mlx-omni-server

> Learn to use the Whisper speech-to-text endpoint in mlx-omni-server. Transcribe audio files to JSON, SRT, VTT, and plain text formats with our easy API.

- Repository: [madroid/mlx-omni-server](https://github.com/madroidmaq/mlx-omni-server)
- Tags: how-to-guide
- Published: 2026-03-06

---

**The mlx-omni-server exposes Whisper speech-to-text capabilities via `POST /audio/transcriptions` and `POST /v1/audio/transcriptions`, accepting audio files and returning transcriptions in multiple formats including JSON, SRT, VTT, and plain text.**

The mlx-omni-server project provides OpenAI-compatible API endpoints for MLX-based models. Its speech-to-text implementation leverages Apple's MLX framework to run Whisper inference locally, offering a privacy-focused alternative to cloud-based transcription services while maintaining API compatibility with OpenAI's audio transcriptions specification.

## Endpoint Routes and Architecture

The STT service is implemented in [`src/mlx_omni_server/stt/stt.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/stt.py) and mounted via [`src/mlx_omni_server/routers.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py). Two identical endpoints handle incoming requests:

- `POST /audio/transcriptions`
- `POST /v1/audio/transcriptions`

Both routes utilize the **`STTService`** class defined in [`src/mlx_omni_server/stt/whisper_model.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/whisper_model.py), which orchestrates temporary file handling and invokes the underlying Whisper model.

## Request Validation and Parameters

Incoming requests are validated through the **`STTRequestForm`** class located in [`src/mlx_omni_server/stt/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/schema.py). This form performs several critical validations before processing begins:

- **File type verification** for uploaded audio files
- **Temperature range constraints** (typically 0.0 to 1.0)
- **Language code validation** (e.g., `en`, `es`, `fr`)
- **Compatibility checks** between `timestamp_granularities` and `response_format`

The form normalizes the `timestamp_granularities[]` form field into a list of **`TimestampGranularity`** enum values. Valid options include `segment` for segment-level timestamps and `word` for word-level timestamps.

## Supported Response Formats

The endpoint supports five distinct output formats controlled by the `response_format` parameter:

| Format | FastAPI Response Type | Content Description |
|--------|----------------------|---------------------|
| `json` | `JSONResponse` | Simple JSON object with `{"text": "transcription"}` |
| `verbose_json` | `JSONResponse` | Complete Whisper output including segments, words, duration, and metadata |
| `text` | `PlainTextResponse` | Raw transcription text only |
| `srt` | `Response` with headers | SubRip subtitle format with `Content-Disposition` for download |
| `vtt` | `Response` with headers | WebVTT subtitle format with `Content-Disposition` for download |

**Important:** Word-level timestamps (`timestamp_granularities[]=word`) require `response_format=verbose_json`. The `STTRequestForm` enforces this compatibility constraint during validation.

## Implementation Workflow

When processing a request, the `STTService` executes the following workflow:

1. Saves the uploaded audio file to a temporary location on disk
2. Calls **`mlx_whisper.transcribe`** with the specified model name, language, temperature, optional prompt, and timestamp granularity flags
3. Formats the raw inference results according to the requested `response_format`
4. Returns the appropriate FastAPI response type based on the format selection

## Practical Code Examples

### cURL Request

```bash
curl -X POST "http://localhost:8000/audio/transcriptions" \
  -H "Accept: application/json" \
  -F "file=@/path/to/audio.wav" \
  -F "model=mlx/whisper-base" \
  -F "language=en" \
  -F "response_format=json" \
  -F "temperature=0.0" \
  -F "timestamp_granularities[]=segment"

```

### Python with requests

```python
import requests

url = "http://localhost:8000/v1/audio/transcriptions"
files = {"file": open("speech.mp3", "rb")}
data = {
    "model": "mlx/whisper-base",
    "language": "en",
    "response_format": "verbose_json",
    "temperature": "0.0",
    "timestamp_granularities[]": "word",  # Requires verbose_json

}

resp = requests.post(url, files=files, data=data)
result = resp.json()  # Contains full segments and word-level timestamps

```

### Python with httpx (Async)

```python
import httpx

async def transcribe():
    async with httpx.AsyncClient(base_url="http://localhost:8000") as client:
        files = {"file": ("audio.wav", open("audio.wav", "rb"), "audio/wav")}
        data = {
            "model": "mlx/whisper-large",
            "response_format": "srt",
            "timestamp_granularities[]": "segment",
        }
        r = await client.post("/audio/transcriptions", files=files, data=data)
        r.raise_for_status()
        print(r.text)  # SRT subtitle content

# asyncio.run(transcribe())

```

## Verbose JSON Response Structure

When requesting `verbose_json`, the endpoint returns a comprehensive payload including timing metadata:

```json
{
  "task": "transcribe",
  "language": "en",
  "duration": 12.34,
  "text": "Hello world, this is a test.",
  "words": [
    {"word": "Hello", "start": 0.0, "end": 0.5},
    {"word": "world", "start": 0.5, "end": 0.9}
  ],
  "segments": [
    {
      "id": 0,
      "seek": 0,
      "start": 0.0,
      "end": 2.1,
      "text": "Hello world",
      "tokens": [50364, 292, 1234],
      "temperature": 0.0,
      "avg_logprob": -0.02,
      "compression_ratio": 1.2,
      "no_speech_prob": 0.01
    }
  ]
}

```

## Summary

- The mlx-omni-server provides OpenAI-compatible speech-to-text endpoints at `/audio/transcriptions` and `/v1/audio/transcriptions`
- Request validation occurs through `STTRequestForm` in [`src/mlx_omni_server/stt/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/schema.py), ensuring parameter compatibility
- The service supports five response formats: `json`, `verbose_json`, `text`, `srt`, and `vtt`
- Word-level timestamps require `verbose_json` format with `timestamp_granularities[]=word`
- Underlying inference uses `mlx_whisper.transcribe` as implemented in [`src/mlx_omni_server/stt/whisper_model.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/whisper_model.py)

## Frequently Asked Questions

### What audio file formats does the Whisper endpoint accept?

The endpoint accepts standard audio formats supported by the underlying MLX Whisper implementation. The `STTRequestForm` in [`src/mlx_omni_server/stt/schema.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/schema.py) validates file types upon upload, ensuring compatibility before the audio reaches the `mlx_whisper.transcribe` call in [`src/mlx_omni_server/stt/whisper_model.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/whisper_model.py).

### How do I enable word-level timestamps in the transcription?

Set `response_format` to `verbose_json` and include `timestamp_granularities[]=word` in your request parameters. The validation logic in `STTRequestForm` enforces that word-level granularities are only requested with verbose JSON format, as this is the only format that includes the detailed `words` array in the response structure.

### Can I use this endpoint as a drop-in replacement for OpenAI's Whisper API?

Yes, the endpoint routes and request parameters mirror the OpenAI Audio Transcriptions API specification. However, you must use MLX-specific model names with the `mlx/` prefix (e.g., `mlx/whisper-base` or `mlx/whisper-large`) as defined in the `STTService` implementation in [`src/mlx_omni_server/stt/whisper_model.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/whisper_model.py).

### Where is the transcription logic implemented in the source code?

The core transcription logic resides in [`src/mlx_omni_server/stt/whisper_model.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/whisper_model.py), specifically within the `STTService` class that handles file storage and calls `mlx_whisper.transcribe`. API routing and response formatting are defined in [`src/mlx_omni_server/stt/stt.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/stt/stt.py), with the router registered in the main application via [`src/mlx_omni_server/routers.py`](https://github.com/madroidmaq/mlx-omni-server/blob/main/src/mlx_omni_server/routers.py).