How to Integrate Whisper for Speech-to-Text Transcription in Voicebox

Voicebox provides a plug-and-play Whisper integration that automatically selects the optimal inference engine (PyTorch or MLX), handles on-demand model downloading, and exposes transcription capabilities through a FastAPI endpoint or direct Python API calls.

Voicebox is an open-source voice processing framework by jamiepine that abstracts OpenAI's Whisper models behind a unified backend interface. When integrating Whisper for transcription with Voicebox, you leverage an asynchronous architecture that caches models in memory, supports concurrent requests, and automatically detects whether to use MLX (Apple Silicon) or PyTorch (CUDA/CPU) based on the host platform.

Architecture Overview

The transcription pipeline follows a clean abstraction pattern defined in backend/backends/__init__.py. When you initiate a transcription request, Voicebox routes the audio through a platform-specific backend while maintaining a consistent interface.

Backend Selection and Protocol

The get_stt_backend() function (line 85 in backend/backends/__init__.py) detects your platform via get_backend_type() and returns either an MLXSTTBackend or PyTorchSTTBackend instance. Both classes implement the STTBackend protocol, ensuring the rest of the codebase remains agnostic to the underlying inference library.

Model Lifecycle Management

Voicebox handles Whisper models through a lazy-loading lifecycle:

  1. is_loaded(): Checks if the model currently resides in memory
  2. _is_model_cached(): Verifies whether the model weights exist in the local Hugging Face cache
  3. load_model_async(): Downloads the model on-demand if missing, reporting progress via model_load_progress
  4. unload_model(): Frees GPU/CPU memory when transcription tasks complete

This approach ensures that subsequent transcription requests use pre-loaded models without the latency of repeated initialization.

HTTP API Integration

The /transcribe endpoint in backend/routes/transcription.py (lines 19-78) provides the primary interface for speech-to-text conversion. It handles file uploads, audio preprocessing, and backend orchestration asynchronously.

Basic cURL Transcription Request

Send an audio file to the FastAPI endpoint with optional parameters for language and model size:

curl -X POST "http://localhost:8000/transcribe" \
     -F "file=@speech.wav" \
     -F "language=en" \
     -F "model=small"

Handling Background Model Downloads

If the requested Whisper model is not cached locally, Voicebox returns HTTP 202 Accepted with a progress indicator:

{
  "message": "Whisper model small is being downloaded. Please wait and try again.",
  "model_name": "whisper-small",
  "downloading": true
}

Retry the request after the download completes to receive the transcription result.

Python HTTP Client Implementation

For Python applications, use httpx to handle multipart uploads and response parsing:

import httpx

def transcribe_file(path: str, language: str = "en", model: str | None = None):
    with open(path, "rb") as fp:
        files = {"file": (path, fp, "audio/wav")}
        data = {"language": language}
        if model:
            data["model"] = model

        resp = httpx.post(
            "http://localhost:8000/transcribe", 
            files=files, 
            data=data
        )
        resp.raise_for_status()
        return resp.json()["text"]

print(transcribe_file("speech.wav", language="en", model="base"))

Direct Backend Usage

For applications requiring tighter integration, import the transcription service directly from backend/services/transcribe.py and bypass the HTTP layer.

Async Transcription in Python

The get_whisper_model() function returns the active STT backend instance. Use load_model_async() to ensure the model is ready, then call transcribe():

import asyncio
from voicebox.backend.services import transcribe

async def transcribe_with_whisper(audio_path: str, model: str = "base"):
    whisper = transcribe.get_whisper_model()
    await whisper.load_model_async(model)
    
    text = await whisper.transcribe(
        audio_path, 
        language="en", 
        model_size=model
    )
    return text

if __name__ == "__main__":
    result = asyncio.run(transcribe_with_whisper("speech.wav", "small"))
    print(result)

Audio Preprocessing Pipeline

Before inference, Voicebox processes raw audio through load_audio() in backend/utils/audio.py (lines 47-64). This utility loads the audio file and resamples it to 16 kHz, which is required by Whisper's input specifications. The function handles temporary file management and normalization automatically.

Model Management and Optimization

Memory Optimization

To free system resources after heavy transcription workloads, explicitly unload the model from memory:

from voicebox.backend.services import transcribe

transcribe.unload_whisper_model()

This calls the backend-specific unload_model() method, clearing GPU VRAM or system RAM depending on your platform.

Thread Pool Execution

Both MLXSTTBackend.transcribe() (lines 33-66) and PyTorchSTTBackend.transcribe() (lines 23-63) execute inference inside thread pools to prevent blocking the async event loop. This design allows the FastAPI server to handle concurrent requests while Whisper processes audio in parallel background threads.

Key Source Files and Implementation Details

Understanding these critical files helps when extending Voicebox's transcription capabilities:

Summary

Integrating Whisper for speech-to-text in Voicebox leverages a robust abstraction layer that handles platform detection, model caching, and asynchronous inference. Key implementation points include:

  • Automatic backend selection via get_stt_backend() chooses MLX for Apple Silicon or PyTorch for other platforms
  • Lazy model loading with load_model_async() downloads Whisper weights on first use and caches them for subsequent requests
  • Non-blocking execution using thread pools ensures the FastAPI event loop remains responsive under load
  • HTTP 202 handling indicates when models are downloading, allowing clients to poll for readiness
  • Direct Python API access through backend/services/transcribe.py enables integration without HTTP overhead

Frequently Asked Questions

How does Voicebox choose between MLX and PyTorch backends?

Voicebox detects your platform automatically in backend/backends/__init__.py via get_backend_type(). If you are running on Apple Silicon with MLX support, it instantiates MLXSTTBackend; otherwise, it falls back to PyTorchSTTBackend for CUDA or CPU inference. Both implement the same STTBackend protocol, ensuring identical behavior regardless of the underlying engine.

What audio formats does the Voicebox transcription endpoint accept?

The /transcribe endpoint streams uploaded files to temporary WAV storage before processing. The load_audio() utility in backend/utils/audio.py handles format conversion and resamples all input to 16 kHz mono, which is the required sample rate for Whisper models. Standard WAV, MP3, and other common formats supported by your Python audio libraries will process correctly.

Can I run transcription without downloading models repeatedly?

Yes. Voicebox implements model caching through is_loaded() and _is_model_cached() checks. Once a model downloads to the Hugging Face cache and loads into memory, subsequent transcription requests reuse the existing instance. Call unload_whisper_model() only when you need to free memory for other operations, as the cached model provides the lowest latency for repeated transcriptions.

How do I handle transcription for multiple languages?

Pass the language parameter in your API request or Python function call. The Whisper backend accepts standard language codes (e.g., "en" for English, "es" for Spanish) and passes them to the underlying Whisper model. Both the HTTP endpoint and direct backend usage support real-time language specification without requiring model reloads, as Whisper models are multilingual by design.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →