# Configuring Voicebox for CUDA GPU Backend: Complete Setup Guide

> Easily configure Voicebox for CUDA GPU backend. Follow our guide to set up your NVIDIA GPU, drivers, and trigger the automated download for seamless integration.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: how-to-guide
- Published: 2026-04-14

---

**You can configure Voicebox for CUDA by ensuring you have an NVIDIA GPU with compatible drivers, then triggering the automated download via the `/backend/download-cuda` endpoint, which installs the `voicebox-server-cuda` binary and NVIDIA runtime libraries to your local data directory.**

The jamiepine/voicebox repository provides a high-performance speech synthesis engine that leverages NVIDIA GPUs through a dedicated CUDA backend variant. Configuring Voicebox for CUDA GPU backend acceleration requires downloading platform-specific PyInstaller-packed executables and CUDA runtime libraries, which the application manages automatically through its service layer in [`backend/services/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/cuda.py).

## Prerequisites for CUDA Support

Before configuring the CUDA backend, verify your environment meets these requirements:

- **NVIDIA GPU**: A compatible NVIDIA graphics card with sufficient VRAM for text-to-speech inference.
- **CUDA-compatible drivers**: Functional NVIDIA drivers that PyTorch can detect (`torch.cuda.is_available()` must return `True`).
- **Voicebox server running**: The backend service must be active (default `localhost:8000`) to expose the CUDA management API.

## Understanding the CUDA Backend Architecture

The Voicebox CUDA implementation consists of two primary components managed through [`backend/services/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/cuda.py):

**Server Core**: A PyInstaller-packed executable (`voicebox-server-cuda`) that handles TTS inference on the GPU.

**CUDA Runtime Libraries**: A collection of NVIDIA-specific `.so` or `.dll` files (version `cu128-v1` as defined in `CUDA_LIBS_VERSION`) required by PyTorch's CUDA support.

Key source files governing this architecture include:

- **[`backend/services/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/cuda.py)**: Core service handling download, extraction, version checks, and cleanup. Contains functions like `get_cuda_dir()`, `download_cuda_binary()`, and `check_and_update_cuda_binary()`.
- **[`backend/routes/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/cuda.py)**: REST API endpoints for backend status checks and download triggers.
- **[`backend/server.py`](https://github.com/jamiepine/voicebox/blob/main/backend/server.py)** (lines 232-238): Detects the CUDA binary name and automatically sets the environment variable `VOICEBOX_BACKEND_VARIANT=cuda`.
- **[`scripts/package_cuda.py`](https://github.com/jamiepine/voicebox/blob/main/scripts/package_cuda.py)**: Build script that creates the release archives (`voicebox-server-cuda.tar.gz` and `cuda-libs-*.tar.gz`).

## Step-by-Step Configuration Process

Follow these steps to configure Voicebox for the CUDA GPU backend:

1. **Start the Voicebox server** and ensure it is accessible at `http://localhost:8000`.

2. **Check current status** to determine if the CUDA backend is already installed:

   ```bash
   curl http://localhost:8000/backend/cuda-status
   ```

3. **Trigger the download** if the backend is not present:

   ```bash
   curl -X POST http://localhost:8000/backend/download-cuda
   ```

   The server responds with:
   
   ```json
   { "message": "CUDA backend download started", "progress_key": "cuda-backend" }
   ```

4. **Monitor download progress** (optional) using the progress endpoint:

   ```bash
   curl http://localhost:8000/backend/cuda-progress
   ```

   This streams Server-Sent Events showing real-time download status.

5. **Verify installation** by checking the status endpoint again:

   ```bash
   curl http://localhost:8000/backend/cuda-status
   ```

   A successful installation returns:
   
   ```json
   {
     "available": true,
     "active": false,
     "binary_path": "/home/user/.local/share/voicebox/backends/cuda/voicebox-server-cuda",
     "cuda_libs_version": "cu128-v1",
     "downloading": false,
     "download_progress": null
   }
   ```

6. **Activate the backend** by starting a voice generation request. The server binary automatically sets `VOICEBOX_BACKEND_VARIANT=cuda` when spawning the CUDA-specific executable, routing all inference to the GPU.

## REST API Endpoints for CUDA Management

The [`backend/routes/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/cuda.py) module exposes these endpoints for backend management:

- **`GET /backend/cuda-status`**: Returns JSON with availability, active state, binary path, and current download progress.
- **`POST /backend/download-cuda`**: Initiates background download of the server binary and CUDA libraries via `download_cuda_binary()`.
- **`DELETE /backend/cuda`**: Removes the entire CUDA backend directory using `delete_cuda_binary()`; fails if the backend is currently active.
- **`GET /backend/cuda-progress`**: Streams SSE events for download progress updates consumed by the web UI.

## Programmatic Configuration with Python

Automate the setup process using the Voicebox REST API:

```python
import httpx
import asyncio

API = "http://localhost:8000"

async def ensure_cuda():
    async with httpx.AsyncClient() as client:
        # 1. Check current status

        status = (await client.get(f"{API}/backend/cuda-status")).json()
        if status["available"]:
            print("CUDA backend already installed.")
            return

        # 2. Trigger download

        resp = await client.post(f"{API}/backend/download-cuda")
        print(resp.json()["message"])

        # 3. Poll until complete

        while True:
            progress = (await client.get(f"{API}/backend/cuda-status")).json()
            if not progress.get("downloading"):
                break
            print(f"Downloading… {progress['download_progress']}")
            await asyncio.sleep(1)

        print("CUDA backend ready!")

asyncio.run(ensure_cuda())

```

## Backend Maintenance and Updates

Voicebox maintains the CUDA backend through automatic and manual update mechanisms:

**Auto-Update**: On server startup, `check_and_update_cuda_binary()` in [`backend/services/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/cuda.py) verifies the installed version against the expected `CUDA_LIBS_VERSION` and automatically downloads updates when the version changes.

**Manual Reinstallation**: To force a fresh installation:

```bash
curl -X DELETE http://localhost:8000/backend/cuda
curl -X POST http://localhost:8000/backend/download-cuda

```

The `delete_cuda_binary()` function removes the entire `<data_dir>/backends/cuda/` directory, while integrity checks during download verify SHA-256 checksums before extraction.

## Summary

Configuring Voicebox for the CUDA GPU backend requires:

- **Hardware verification**: Ensure NVIDIA GPU and drivers are present and detected by PyTorch.
- **Automated download**: Use the `/backend/download-cuda` endpoint to fetch the `voicebox-server-cuda` binary and `cu128-v1` runtime libraries.
- **Directory management**: Binaries install to `<data_dir>/backends/cuda/` as managed by `get_cuda_dir()` in [`backend/services/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/services/cuda.py).
- **Environment activation**: The system sets `VOICEBOX_BACKEND_VARIANT=cuda` automatically when launching the CUDA server binary detected in [`backend/server.py`](https://github.com/jamiepine/voicebox/blob/main/backend/server.py).
- **API-based management**: Monitor and control the backend via REST endpoints defined in [`backend/routes/cuda.py`](https://github.com/jamiepine/voicebox/blob/main/backend/routes/cuda.py).

## Frequently Asked Questions

### What NVIDIA hardware is required to run Voicebox with CUDA?

Voicebox requires a CUDA-compatible NVIDIA GPU with sufficient VRAM to load the text-to-speech models, plus NVIDIA drivers that PyTorch recognizes. You can verify compatibility by checking if `torch.cuda.is_available()` returns `True` in your Python environment before starting the Voicebox server.

### How can I verify that Voicebox is actually using the GPU for inference?

Check the server logs for the environment variable `VOICEBOX_BACKEND_VARIANT=cuda`, which is set automatically in [`backend/server.py`](https://github.com/jamiepine/voicebox/blob/main/backend/server.py) (lines 232-238) when the CUDA binary is active. Additionally, query `GET /backend/cuda-status` and confirm `"active": true` in the response JSON.

### What happens when Voicebox releases a new CUDA backend version?

When the `CUDA_LIBS_VERSION` constant changes (e.g., from "cu128-v1" to a newer release), the `check_and_update_cuda_binary()` function automatically detects the mismatch on server startup and downloads the updated archives. You can also manually trigger an update by deleting the backend via `DELETE /backend/cuda` and reinstalling.

### Why does the CUDA backend download fail or hang?

Download failures typically occur due to network restrictions, insufficient disk space in the data directory, or conflicts when the backend is currently active. Ensure the server has write permissions to `<data_dir>/backends/cuda/`, verify no generation jobs are running before deletion, and check that your system can reach GitHub releases for the `voicebox-server-cuda` archives.