# Configuring Voicebox for the MLX GPU Backend on Apple Silicon

> Configure Voicebox for MLX GPU backend on Apple Silicon. Install dependencies and run Voicebox to enable Metal GPU acceleration for 4-5x faster inference than CPU.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: how-to-guide
- Published: 2026-04-14

---

**Install the MLX-specific dependencies from [`backend/requirements-mlx.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements-mlx.txt) and run Voicebox on Apple Silicon to automatically activate the Metal GPU backend, which delivers 4-5× faster inference than CPU-based alternatives.**

Voicebox is an open-source text-to-speech and speech-to-text framework that dynamically selects its inference engine based on the host platform. Configuring Voicebox for the MLX GPU backend requires no manual code modifications—simply ensure you have the correct Python dependencies installed and run on compatible Apple Silicon hardware.

## Installing MLX Dependencies

The MLX backend requires specific packages listed in **[`backend/requirements-mlx.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements-mlx.txt)**. These dependencies enable Metal Performance Shaders acceleration on macOS.

Install the requirements using pip:

```bash
pip install -r backend/requirements-mlx.txt

```

This installs **`mlx>=0.30.0`** and **`mlx-audio>=0.3.1`**, which provide the core tensor operations and audio processing libraries for Apple Silicon.

## Platform Detection and Backend Selection

Voicebox determines which inference engine to use via **[`backend/utils/platform_detect.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/platform_detect.py)**. The `get_backend_type()` function attempts to import `mlx.core` and returns `"mlx"` when running on Apple Silicon with the package available.

You can verify the detected backend programmatically:

```python
from backend.utils.platform_detect import get_backend_type

print("Detected backend:", get_backend_type())  # Outputs: "mlx"

```

To force MLX detection even on non-Apple-Silicon machines for testing purposes, set the environment variable:

```bash
export VOICEBOX_FORCE_MLX=1

```

## Backend Factory and MLX Implementation

When `"mlx"` is detected, the factory in **[`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py)** instantiates **`MLXTTSBackend`** for text-to-speech and **`MLXSTTBackend`** for speech-to-text. These classes are defined in **[`backend/backends/mlx_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/mlx_backend.py)** and wrap the `mlx_audio` library while maintaining API compatibility with the PyTorch counterparts.

The backend factory also handles model identifier selection. For the Qwen-TTS models, it automatically selects:
- **MLX**: `mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16`
- **PyTorch**: `Qwen/Qwen3-TTS-12Hz-1.7B-Base`

This ensures you receive the correctly quantized model weights for your compute backend.

## Offline Caching and Model Storage

Voicebox patches Hugging Face download behavior via **[`backend/utils/hf_offline_patch.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/hf_offline_patch.py)**. This module forces offline mode when models are already cached and creates symlinks so that MLX-specific Qwen repositories can reuse cached original Qwen model configurations.

To customize the model cache location, set the **`VOICEBOX_MODELS_DIR`** environment variable. This updates both the Voicebox internal cache and the **`HF_HUB_CACHE`** variable used by the Hugging Face Hub library:

```bash
export VOICEBOX_MODELS_DIR="$HOME/.cache/voicebox_models"

```

The configuration logic resides in **[`backend/config.py`](https://github.com/jamiepine/voicebox/blob/main/backend/config.py)**, which reads these environment variables at startup.

## Verifying MLX Activation

When the server initializes, it logs the active backend. You should see a line containing `Backend: MLX` in the console output. The health check endpoint defined in **[`backend/routes/health.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/health.py)** also reports the current backend type, confirming that audio inference will utilize the Metal GPU.

To synthesize speech using the verified MLX backend:

```python
import asyncio
from backend.backends import get_tts_backend

async def synthesize():
    tts = get_tts_backend()  # Returns MLXTTSBackend instance

    await tts.load_model_async("1.7B")
    
    voice_prompt, _ = await tts.create_voice_prompt(
        audio_path="reference.wav",
        reference_text="Hello, this is my voice sample."
    )
    
    audio, sr = await tts.generate(
        text="Voicebox is now running on MLX GPU acceleration.",
        voice_prompt=voice_prompt,
        language="en"
    )
    
    import soundfile as sf
    sf.write("output.wav", audio, sr)

asyncio.run(synthesize())

```

## Complete Configuration Workflow

Follow these steps to configure Voicebox for MLX GPU inference:

1. **Install dependencies**:
   ```bash
   pip install -r backend/requirements-mlx.txt
   ```

2. **(Optional) Set custom cache directory**:
   ```bash
   export VOICEBOX_MODELS_DIR="$HOME/.cache/voicebox_models"
   ```

3. **(Optional) Force MLX on unsupported hardware**:
   ```bash
   export VOICEBOX_FORCE_MLX=1
   ```

4. **Start the server**:
   ```bash
   just dev  # or: python -m backend.main

   ```

   
   Confirm the log shows `Backend: MLX` to verify Metal GPU acceleration is active.

## Summary

- **Install MLX packages** from [`backend/requirements-mlx.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements-mlx.txt) (`mlx>=0.30.0` and `mlx-audio>=0.3.1`).
- **Automatic detection** occurs in [`backend/utils/platform_detect.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/platform_detect.py) when running on Apple Silicon.
- **Backend instantiation** happens via the factory in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py), selecting `MLXTTSBackend` from [`backend/backends/mlx_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/mlx_backend.py).
- **Model management** uses MLX-optimized weights (`mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16`) and supports offline caching via [`backend/utils/hf_offline_patch.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/hf_offline_patch.py).
- **Verify configuration** by checking server logs for "Backend: MLX" or querying the health endpoint in [`backend/routes/health.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/health.py).

## Frequently Asked Questions

### Do I need to modify any code to enable MLX acceleration?

No code modifications are necessary. Simply install the dependencies from [`backend/requirements-mlx.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements-mlx.txt) and run Voicebox on Apple Silicon. The `get_backend_type()` function in [`backend/utils/platform_detect.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/platform_detect.py) automatically detects the platform and imports `mlx.core` to activate the Metal GPU backend.

### Can I use the MLX backend on Intel Macs or Linux machines?

No, the MLX backend requires Apple Silicon (M1/M2/M3/M4 chips). However, you can force the detection logic for testing purposes by setting `VOICEBOX_FORCE_MLX=1` before starting the server, though this will fail if the underlying `mlx` package cannot import the required Metal libraries.

### How does Voicebox handle model downloads for the MLX backend?

Voicebox downloads Qwen-TTS models from the `mlx-community` organization on Hugging Face (specifically `mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16`). The [`backend/utils/hf_offline_patch.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/hf_offline_patch.py) module ensures that cached models are reused and creates necessary symlinks for configuration compatibility, allowing offline operation once models are downloaded.

### Where is the model cache stored when using MLX?

By default, models are stored in the standard Hugging Face cache location. You can override this by setting the `VOICEBOX_MODELS_DIR` environment variable, which updates the cache path in [`backend/config.py`](https://github.com/jamiepine/voicebox/blob/main/backend/config.py) and also sets `HF_HUB_CACHE` for all Hugging Face downloads within the application.