Configuring Voicebox for CUDA GPU Backend: Complete Setup Guide
You can configure Voicebox for CUDA by ensuring you have an NVIDIA GPU with compatible drivers, then triggering the automated download via the /backend/download-cuda endpoint, which installs the voicebox-server-cuda binary and NVIDIA runtime libraries to your local data directory.
The jamiepine/voicebox repository provides a high-performance speech synthesis engine that leverages NVIDIA GPUs through a dedicated CUDA backend variant. Configuring Voicebox for CUDA GPU backend acceleration requires downloading platform-specific PyInstaller-packed executables and CUDA runtime libraries, which the application manages automatically through its service layer in backend/services/cuda.py.
Prerequisites for CUDA Support
Before configuring the CUDA backend, verify your environment meets these requirements:
- NVIDIA GPU: A compatible NVIDIA graphics card with sufficient VRAM for text-to-speech inference.
- CUDA-compatible drivers: Functional NVIDIA drivers that PyTorch can detect (
torch.cuda.is_available()must returnTrue). - Voicebox server running: The backend service must be active (default
localhost:8000) to expose the CUDA management API.
Understanding the CUDA Backend Architecture
The Voicebox CUDA implementation consists of two primary components managed through backend/services/cuda.py:
Server Core: A PyInstaller-packed executable (voicebox-server-cuda) that handles TTS inference on the GPU.
CUDA Runtime Libraries: A collection of NVIDIA-specific .so or .dll files (version cu128-v1 as defined in CUDA_LIBS_VERSION) required by PyTorch's CUDA support.
Key source files governing this architecture include:
backend/services/cuda.py: Core service handling download, extraction, version checks, and cleanup. Contains functions likeget_cuda_dir(),download_cuda_binary(), andcheck_and_update_cuda_binary().backend/routes/cuda.py: REST API endpoints for backend status checks and download triggers.backend/server.py(lines 232-238): Detects the CUDA binary name and automatically sets the environment variableVOICEBOX_BACKEND_VARIANT=cuda.scripts/package_cuda.py: Build script that creates the release archives (voicebox-server-cuda.tar.gzandcuda-libs-*.tar.gz).
Step-by-Step Configuration Process
Follow these steps to configure Voicebox for the CUDA GPU backend:
-
Start the Voicebox server and ensure it is accessible at
http://localhost:8000. -
Check current status to determine if the CUDA backend is already installed:
curl http://localhost:8000/backend/cuda-status -
Trigger the download if the backend is not present:
curl -X POST http://localhost:8000/backend/download-cudaThe server responds with:
{ "message": "CUDA backend download started", "progress_key": "cuda-backend" } -
Monitor download progress (optional) using the progress endpoint:
curl http://localhost:8000/backend/cuda-progressThis streams Server-Sent Events showing real-time download status.
-
Verify installation by checking the status endpoint again:
curl http://localhost:8000/backend/cuda-statusA successful installation returns:
{ "available": true, "active": false, "binary_path": "/home/user/.local/share/voicebox/backends/cuda/voicebox-server-cuda", "cuda_libs_version": "cu128-v1", "downloading": false, "download_progress": null } -
Activate the backend by starting a voice generation request. The server binary automatically sets
VOICEBOX_BACKEND_VARIANT=cudawhen spawning the CUDA-specific executable, routing all inference to the GPU.
REST API Endpoints for CUDA Management
The backend/routes/cuda.py module exposes these endpoints for backend management:
GET /backend/cuda-status: Returns JSON with availability, active state, binary path, and current download progress.POST /backend/download-cuda: Initiates background download of the server binary and CUDA libraries viadownload_cuda_binary().DELETE /backend/cuda: Removes the entire CUDA backend directory usingdelete_cuda_binary(); fails if the backend is currently active.GET /backend/cuda-progress: Streams SSE events for download progress updates consumed by the web UI.
Programmatic Configuration with Python
Automate the setup process using the Voicebox REST API:
import httpx
import asyncio
API = "http://localhost:8000"
async def ensure_cuda():
async with httpx.AsyncClient() as client:
# 1. Check current status
status = (await client.get(f"{API}/backend/cuda-status")).json()
if status["available"]:
print("CUDA backend already installed.")
return
# 2. Trigger download
resp = await client.post(f"{API}/backend/download-cuda")
print(resp.json()["message"])
# 3. Poll until complete
while True:
progress = (await client.get(f"{API}/backend/cuda-status")).json()
if not progress.get("downloading"):
break
print(f"Downloading… {progress['download_progress']}")
await asyncio.sleep(1)
print("CUDA backend ready!")
asyncio.run(ensure_cuda())
Backend Maintenance and Updates
Voicebox maintains the CUDA backend through automatic and manual update mechanisms:
Auto-Update: On server startup, check_and_update_cuda_binary() in backend/services/cuda.py verifies the installed version against the expected CUDA_LIBS_VERSION and automatically downloads updates when the version changes.
Manual Reinstallation: To force a fresh installation:
curl -X DELETE http://localhost:8000/backend/cuda
curl -X POST http://localhost:8000/backend/download-cuda
The delete_cuda_binary() function removes the entire <data_dir>/backends/cuda/ directory, while integrity checks during download verify SHA-256 checksums before extraction.
Summary
Configuring Voicebox for the CUDA GPU backend requires:
- Hardware verification: Ensure NVIDIA GPU and drivers are present and detected by PyTorch.
- Automated download: Use the
/backend/download-cudaendpoint to fetch thevoicebox-server-cudabinary andcu128-v1runtime libraries. - Directory management: Binaries install to
<data_dir>/backends/cuda/as managed byget_cuda_dir()inbackend/services/cuda.py. - Environment activation: The system sets
VOICEBOX_BACKEND_VARIANT=cudaautomatically when launching the CUDA server binary detected inbackend/server.py. - API-based management: Monitor and control the backend via REST endpoints defined in
backend/routes/cuda.py.
Frequently Asked Questions
What NVIDIA hardware is required to run Voicebox with CUDA?
Voicebox requires a CUDA-compatible NVIDIA GPU with sufficient VRAM to load the text-to-speech models, plus NVIDIA drivers that PyTorch recognizes. You can verify compatibility by checking if torch.cuda.is_available() returns True in your Python environment before starting the Voicebox server.
How can I verify that Voicebox is actually using the GPU for inference?
Check the server logs for the environment variable VOICEBOX_BACKEND_VARIANT=cuda, which is set automatically in backend/server.py (lines 232-238) when the CUDA binary is active. Additionally, query GET /backend/cuda-status and confirm "active": true in the response JSON.
What happens when Voicebox releases a new CUDA backend version?
When the CUDA_LIBS_VERSION constant changes (e.g., from "cu128-v1" to a newer release), the check_and_update_cuda_binary() function automatically detects the mismatch on server startup and downloads the updated archives. You can also manually trigger an update by deleting the backend via DELETE /backend/cuda and reinstalling.
Why does the CUDA backend download fail or hang?
Download failures typically occur due to network restrictions, insufficient disk space in the data directory, or conflicts when the backend is currently active. Ensure the server has write permissions to <data_dir>/backends/cuda/, verify no generation jobs are running before deletion, and check that your system can reach GitHub releases for the voicebox-server-cuda archives.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →