Configuring Voicebox for the MLX GPU Backend on Apple Silicon

Install the MLX-specific dependencies from backend/requirements-mlx.txt and run Voicebox on Apple Silicon to automatically activate the Metal GPU backend, which delivers 4-5× faster inference than CPU-based alternatives.

Voicebox is an open-source text-to-speech and speech-to-text framework that dynamically selects its inference engine based on the host platform. Configuring Voicebox for the MLX GPU backend requires no manual code modifications—simply ensure you have the correct Python dependencies installed and run on compatible Apple Silicon hardware.

Installing MLX Dependencies

The MLX backend requires specific packages listed in backend/requirements-mlx.txt. These dependencies enable Metal Performance Shaders acceleration on macOS.

Install the requirements using pip:

pip install -r backend/requirements-mlx.txt

This installs mlx>=0.30.0 and mlx-audio>=0.3.1, which provide the core tensor operations and audio processing libraries for Apple Silicon.

Platform Detection and Backend Selection

Voicebox determines which inference engine to use via backend/utils/platform_detect.py. The get_backend_type() function attempts to import mlx.core and returns "mlx" when running on Apple Silicon with the package available.

You can verify the detected backend programmatically:

from backend.utils.platform_detect import get_backend_type

print("Detected backend:", get_backend_type())  # Outputs: "mlx"

To force MLX detection even on non-Apple-Silicon machines for testing purposes, set the environment variable:

export VOICEBOX_FORCE_MLX=1

Backend Factory and MLX Implementation

When "mlx" is detected, the factory in backend/backends/__init__.py instantiates MLXTTSBackend for text-to-speech and MLXSTTBackend for speech-to-text. These classes are defined in backend/backends/mlx_backend.py and wrap the mlx_audio library while maintaining API compatibility with the PyTorch counterparts.

The backend factory also handles model identifier selection. For the Qwen-TTS models, it automatically selects:

  • MLX: mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16
  • PyTorch: Qwen/Qwen3-TTS-12Hz-1.7B-Base

This ensures you receive the correctly quantized model weights for your compute backend.

Offline Caching and Model Storage

Voicebox patches Hugging Face download behavior via backend/utils/hf_offline_patch.py. This module forces offline mode when models are already cached and creates symlinks so that MLX-specific Qwen repositories can reuse cached original Qwen model configurations.

To customize the model cache location, set the VOICEBOX_MODELS_DIR environment variable. This updates both the Voicebox internal cache and the HF_HUB_CACHE variable used by the Hugging Face Hub library:

export VOICEBOX_MODELS_DIR="$HOME/.cache/voicebox_models"

The configuration logic resides in backend/config.py, which reads these environment variables at startup.

Verifying MLX Activation

When the server initializes, it logs the active backend. You should see a line containing Backend: MLX in the console output. The health check endpoint defined in backend/routes/health.py also reports the current backend type, confirming that audio inference will utilize the Metal GPU.

To synthesize speech using the verified MLX backend:

import asyncio
from backend.backends import get_tts_backend

async def synthesize():
    tts = get_tts_backend()  # Returns MLXTTSBackend instance

    await tts.load_model_async("1.7B")
    
    voice_prompt, _ = await tts.create_voice_prompt(
        audio_path="reference.wav",
        reference_text="Hello, this is my voice sample."
    )
    
    audio, sr = await tts.generate(
        text="Voicebox is now running on MLX GPU acceleration.",
        voice_prompt=voice_prompt,
        language="en"
    )
    
    import soundfile as sf
    sf.write("output.wav", audio, sr)

asyncio.run(synthesize())

Complete Configuration Workflow

Follow these steps to configure Voicebox for MLX GPU inference:

  1. Install dependencies:

    pip install -r backend/requirements-mlx.txt
  2. (Optional) Set custom cache directory:

    export VOICEBOX_MODELS_DIR="$HOME/.cache/voicebox_models"
  3. (Optional) Force MLX on unsupported hardware:

    export VOICEBOX_FORCE_MLX=1
  4. Start the server:

    just dev  # or: python -m backend.main
    

    Confirm the log shows Backend: MLX to verify Metal GPU acceleration is active.

Summary

Frequently Asked Questions

Do I need to modify any code to enable MLX acceleration?

No code modifications are necessary. Simply install the dependencies from backend/requirements-mlx.txt and run Voicebox on Apple Silicon. The get_backend_type() function in backend/utils/platform_detect.py automatically detects the platform and imports mlx.core to activate the Metal GPU backend.

Can I use the MLX backend on Intel Macs or Linux machines?

No, the MLX backend requires Apple Silicon (M1/M2/M3/M4 chips). However, you can force the detection logic for testing purposes by setting VOICEBOX_FORCE_MLX=1 before starting the server, though this will fail if the underlying mlx package cannot import the required Metal libraries.

How does Voicebox handle model downloads for the MLX backend?

Voicebox downloads Qwen-TTS models from the mlx-community organization on Hugging Face (specifically mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16). The backend/utils/hf_offline_patch.py module ensures that cached models are reused and creates necessary symlinks for configuration compatibility, allowing offline operation once models are downloaded.

Where is the model cache stored when using MLX?

By default, models are stored in the standard Hugging Face cache location. You can override this by setting the VOICEBOX_MODELS_DIR environment variable, which updates the cache path in backend/config.py and also sets HF_HUB_CACHE for all Hugging Face downloads within the application.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →