# How to Integrate Whisper for Speech-to-Text Transcription in Voicebox

> Integrate Whisper for speech-to-text transcription with Voicebox. Enjoy automatic engine selection, on-demand model downloads, and flexible API access for seamless integration.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: how-to-guide
- Published: 2026-04-14

---

**Voicebox provides a plug-and-play Whisper integration that automatically selects the optimal inference engine (PyTorch or MLX), handles on-demand model downloading, and exposes transcription capabilities through a FastAPI endpoint or direct Python API calls.**

Voicebox is an open-source voice processing framework by `jamiepine` that abstracts OpenAI's Whisper models behind a unified backend interface. When integrating Whisper for transcription with Voicebox, you leverage an asynchronous architecture that caches models in memory, supports concurrent requests, and automatically detects whether to use **MLX** (Apple Silicon) or **PyTorch** (CUDA/CPU) based on the host platform.

## Architecture Overview

The transcription pipeline follows a clean abstraction pattern defined in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py). When you initiate a transcription request, Voicebox routes the audio through a platform-specific backend while maintaining a consistent interface.

### Backend Selection and Protocol

The `get_stt_backend()` function (line 85 in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py)) detects your platform via `get_backend_type()` and returns either an `MLXSTTBackend` or `PyTorchSTTBackend` instance. Both classes implement the **STTBackend** protocol, ensuring the rest of the codebase remains agnostic to the underlying inference library.

- **MLXSTTBackend**: Optimized for Apple Silicon and MLX-compatible Linux systems ([`backend/backends/mlx_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/mlx_backend.py))
- **PyTorchSTTBackground**: Fallback for CUDA and CPU inference ([`backend/backends/pytorch_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/pytorch_backend.py))

### Model Lifecycle Management

Voicebox handles Whisper models through a lazy-loading lifecycle:

1. **`is_loaded()`**: Checks if the model currently resides in memory
2. **`_is_model_cached()`**: Verifies whether the model weights exist in the local Hugging Face cache
3. **`load_model_async()`**: Downloads the model on-demand if missing, reporting progress via `model_load_progress`
4. **`unload_model()`**: Frees GPU/CPU memory when transcription tasks complete

This approach ensures that subsequent transcription requests use pre-loaded models without the latency of repeated initialization.

## HTTP API Integration

The `/transcribe` endpoint in [`backend/routes/transcription.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/transcription.py) (lines 19-78) provides the primary interface for speech-to-text conversion. It handles file uploads, audio preprocessing, and backend orchestration asynchronously.

### Basic cURL Transcription Request

Send an audio file to the FastAPI endpoint with optional parameters for language and model size:

```bash
curl -X POST "http://localhost:8000/transcribe" \
     -F "file=@speech.wav" \
     -F "language=en" \
     -F "model=small"

```

### Handling Background Model Downloads

If the requested Whisper model is not cached locally, Voicebox returns **HTTP 202 Accepted** with a progress indicator:

```json
{
  "message": "Whisper model small is being downloaded. Please wait and try again.",
  "model_name": "whisper-small",
  "downloading": true
}

```

Retry the request after the download completes to receive the transcription result.

### Python HTTP Client Implementation

For Python applications, use `httpx` to handle multipart uploads and response parsing:

```python
import httpx

def transcribe_file(path: str, language: str = "en", model: str | None = None):
    with open(path, "rb") as fp:
        files = {"file": (path, fp, "audio/wav")}
        data = {"language": language}
        if model:
            data["model"] = model

        resp = httpx.post(
            "http://localhost:8000/transcribe", 
            files=files, 
            data=data
        )
        resp.raise_for_status()
        return resp.json()["text"]

print(transcribe_file("speech.wav", language="en", model="base"))

```

## Direct Backend Usage

For applications requiring tighter integration, import the transcription service directly from [`backend/services/transcribe.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/transcribe.py) and bypass the HTTP layer.

### Async Transcription in Python

The `get_whisper_model()` function returns the active STT backend instance. Use `load_model_async()` to ensure the model is ready, then call `transcribe()`:

```python
import asyncio
from voicebox.backend.services import transcribe

async def transcribe_with_whisper(audio_path: str, model: str = "base"):
    whisper = transcribe.get_whisper_model()
    await whisper.load_model_async(model)
    
    text = await whisper.transcribe(
        audio_path, 
        language="en", 
        model_size=model
    )
    return text

if __name__ == "__main__":
    result = asyncio.run(transcribe_with_whisper("speech.wav", "small"))
    print(result)

```

### Audio Preprocessing Pipeline

Before inference, Voicebox processes raw audio through `load_audio()` in [`backend/utils/audio.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/audio.py) (lines 47-64). This utility loads the audio file and resamples it to 16 kHz, which is required by Whisper's input specifications. The function handles temporary file management and normalization automatically.

## Model Management and Optimization

### Memory Optimization

To free system resources after heavy transcription workloads, explicitly unload the model from memory:

```python
from voicebox.backend.services import transcribe

transcribe.unload_whisper_model()

```

This calls the backend-specific `unload_model()` method, clearing GPU VRAM or system RAM depending on your platform.

### Thread Pool Execution

Both `MLXSTTBackend.transcribe()` (lines 33-66) and `PyTorchSTTBackend.transcribe()` (lines 23-63) execute inference inside thread pools to prevent blocking the async event loop. This design allows the FastAPI server to handle concurrent requests while Whisper processes audio in parallel background threads.

## Key Source Files and Implementation Details

Understanding these critical files helps when extending Voicebox's transcription capabilities:

- **[`backend/services/transcribe.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/transcribe.py)**: Entry point providing `get_whisper_model()` and `unload_whisper_model()` helpers
- **[`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py)**: Factory logic at line 85 for `get_stt_backend()` and the `STTBackend` protocol definition
- **[`backend/backends/pytorch_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/pytorch_backend.py)**: PyTorch implementation with `load_model_async()` and `_is_model_cached()` logic
- **[`backend/backends/mlx_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/mlx_backend.py)**: MLX implementation for Apple Silicon acceleration
- **[`backend/routes/transcription.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/transcription.py)**: FastAPI endpoint handling HTTP requests, temporary file storage, and response formatting (lines 19-78)
- **[`backend/utils/audio.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/audio.py)**: Audio loading and 16 kHz resampling utilities (lines 47-64)
- **[`backend/utils/hf_offline_patch.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/hf_offline_patch.py)**: Forces offline mode when models exist in cache, preventing unnecessary network calls

## Summary

Integrating Whisper for speech-to-text in Voicebox leverages a robust abstraction layer that handles platform detection, model caching, and asynchronous inference. Key implementation points include:

- **Automatic backend selection** via `get_stt_backend()` chooses MLX for Apple Silicon or PyTorch for other platforms
- **Lazy model loading** with `load_model_async()` downloads Whisper weights on first use and caches them for subsequent requests
- **Non-blocking execution** using thread pools ensures the FastAPI event loop remains responsive under load
- **HTTP 202 handling** indicates when models are downloading, allowing clients to poll for readiness
- **Direct Python API** access through [`backend/services/transcribe.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/transcribe.py) enables integration without HTTP overhead

## Frequently Asked Questions

### How does Voicebox choose between MLX and PyTorch backends?

Voicebox detects your platform automatically in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) via `get_backend_type()`. If you are running on Apple Silicon with MLX support, it instantiates `MLXSTTBackend`; otherwise, it falls back to `PyTorchSTTBackend` for CUDA or CPU inference. Both implement the same `STTBackend` protocol, ensuring identical behavior regardless of the underlying engine.

### What audio formats does the Voicebox transcription endpoint accept?

The `/transcribe` endpoint streams uploaded files to temporary WAV storage before processing. The `load_audio()` utility in [`backend/utils/audio.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/audio.py) handles format conversion and resamples all input to 16 kHz mono, which is the required sample rate for Whisper models. Standard WAV, MP3, and other common formats supported by your Python audio libraries will process correctly.

### Can I run transcription without downloading models repeatedly?

Yes. Voicebox implements model caching through `is_loaded()` and `_is_model_cached()` checks. Once a model downloads to the Hugging Face cache and loads into memory, subsequent transcription requests reuse the existing instance. Call `unload_whisper_model()` only when you need to free memory for other operations, as the cached model provides the lowest latency for repeated transcriptions.

### How do I handle transcription for multiple languages?

Pass the `language` parameter in your API request or Python function call. The Whisper backend accepts standard language codes (e.g., "en" for English, "es" for Spanish) and passes them to the underlying Whisper model. Both the HTTP endpoint and direct backend usage support real-time language specification without requiring model reloads, as Whisper models are multilingual by design.