Configuring Voicebox for the ROCm GPU Backend

Voicebox automatically detects AMD GPUs via the ROCm stack by inspecting torch.version.hip in backend/app.py, enabling GPU-accelerated inference without code changes when using Docker or native Linux installations.

Voicebox supports AMD GPUs through the ROCm (Radeon Open Compute) platform. This guide explains how to configure the jamiepine/voicebox repository to leverage ROCm for accelerated text-to-speech and transcription workloads, covering both containerized and native deployment strategies.

How Voicebox Detects ROCm GPUs

The detection logic resides in backend/app.py within the _get_gpu_status helper function. This function checks whether PyTorch reports a ROCm-enabled build by inspecting torch.version.hip. If ROCm is present, the function returns a human-readable string like ROCm (Radeon™ RX 6600 XT) that appears in the server UI and determines which device hosts the model.

def _get_gpu_status() -> str:
    backend_type = get_backend_type()
    if torch.cuda.is_available():
        device_name = torch.cuda.get_device_name(0)
        is_rocm = hasattr(torch.version, "hip") and torch.version.hip is not None
        if is_rocm:
            return f"ROCm ({device_name})"
        return f"CUDA ({device_name})"
    ...

Source: backend/app.py L145–L152

When torch.cuda.is_available() returns True and torch.version.hip exists, Voicebox schedules model execution on the ROCm device. Otherwise, it falls back to CPU, CUDA, or other backends (MPS, XPU, etc.).

Deployment Options for ROCm

You have two primary methods to run Voicebox with ROCm support: Docker-based deployment (experimental) or native Linux installation. Both approaches rely on the same runtime detection in backend/app.py, requiring no application code changes once the environment is configured.

The repository provides an experimental ROCm Dockerfile that installs a ROCm-enabled PyTorch wheel and configures required environment variables. According to docs/plans/DOCKER_DEPLOYMENT.md, the image starts from rocm/dev-ubuntu-22.04:6.0 and targets the PyTorch ROCm 6.0 wheel index.

FROM rocm/dev-ubuntu-22.04:6.0

# Install Python + basic deps

RUN apt-get update && apt-get install -y \
    python3.11 python3-pip git ffmpeg && \
    rm -rf /var/lib/apt/lists/*

WORKDIR /app

# Install ROCm-enabled PyTorch

COPY backend/requirements.txt .
RUN pip3 install torch torchvision torchaudio \
    --index-url https://download.pytorch.org/whl/rocm6.0

# Install the rest of Voicebox's Python deps

RUN pip3 install -r requirements.txt
RUN pip3 install git+https://github.com/QwenLM/Qwen3-TTS.git

# ROCm environment overrides (helps newer GPUs)

ENV HSA_OVERRIDE_GFX_VERSION=10.3.0
ENV ROCM_PATH=/opt/rocm

COPY backend/ /app/backend/
EXPOSE 8000
CMD ["uvicorn", "backend.main:app", "--host", "0.0.0.0", "--port", "8000"]

Source: docs/plans/DOCKER_DEPLOYMENT.md L216–L240

To run the container, you must expose the GPU devices to the container runtime. The documentation specifies mounting /dev/kfd and /dev/dri with specific security options:

docker run --device=/dev/kfd --device=/dev/dri \
  --group-add video --ipc=host --cap-add=SYS_PTRACE \
  --security-opt seccomp=unconfined \
  -p 8000:8000 -v voicebox-data:/app/data \
  voicebox:rocm

Source: docs/plans/DOCKER_DEPLOYMENT.md L247–L254

Docker Compose Configuration

For persistent deployments, docs/content/docs/overview/docker.mdx documents the required docker-compose service definition. You must add the device entries and group permissions to the service:

services:
  voicebox:
    build: .
    devices:
      - /dev/kfd
      - /dev/dri
    group_add:
      - video

Source: docs/content/docs/overview/docker.mdx L28–L42

Native Linux Installation

For bare-metal deployments, install the AMD ROCm drivers on your host system, then manually install the ROCm-compatible PyTorch wheel using the same index URL referenced in the Dockerfile (https://download.pytorch.org/whl/rocm6.0). Ensure the HSA_OVERRIDE_GFX_VERSION and ROCM_PATH environment variables are exported in your shell session before starting Voicebox.

Environment Variables for GPU Compatibility

ROCm support in Voicebox includes specific environment overrides to broaden hardware compatibility. As noted in docs/notes/RELEASE_v0.2.0.md, the HSA_OVERRIDE_GFX_VERSION variable allows newer Radeon GPUs not officially listed in ROCm's compatibility matrix to function correctly.

  • HSA_OVERRIDE_GFX_VERSION: Set to 10.3.0 (or appropriate version) to override the graphics target for unsupported GPUs like the Radeon RX 6600 XT.
  • ROCM_PATH: Points to the ROCm installation directory (typically /opt/rocm).

Source: docs/notes/RELEASE_v0.2.0.md L91

Verifying ROCm Detection

Once deployed, verify ROCm detection by checking the Voicebox server logs or UI. The _get_gpu_status function in backend/app.py will report the device name prefixed with "ROCm" if detection succeeds. If the system falls back to CPU, verify that:

  1. The container or host has access to /dev/kfd and /dev/dri
  2. The video group permissions are correctly applied
  3. torch.version.hip returns a valid version string in your Python environment

Summary

  • Automatic Detection: Voicebox detects ROCm GPUs via torch.version.hip in backend/app.py without requiring manual backend selection.
  • Docker Support: Use the experimental ROCm Dockerfile based on rocm/dev-ubuntu-22.04:6.0 with PyTorch's ROCm 6.0 wheel index.
  • Device Access: Expose /dev/kfd and /dev/dri to containers with --group-add video and --security-opt seccomp=unconfined.
  • Compatibility Overrides: Set HSA_OVERRIDE_GFX_VERSION=10.3.0 for newer AMD GPUs not in the official ROCm support matrix.

Frequently Asked Questions

Does Voicebox support ROCm on Windows?

No. The ROCm support is considered experimental and works best on Linux systems. The Docker deployment uses rocm/dev-ubuntu-22.04:6.0 as its base image, and the device paths (/dev/kfd, /dev/dri) are Linux-specific kernel interfaces.

Which AMD GPUs work with Voicebox and ROCm?

Voicebox relies on PyTorch's ROCm support. While official compatibility varies by ROCm version, the HSA_OVERRIDE_GFX_VERSION environment variable enables support for newer consumer GPUs like the Radeon RX 6600 XT that may not appear in AMD's official compatibility matrix.

Do I need to modify code to enable ROCm?

No. Once the environment is correctly configured with ROCm drivers and the appropriate PyTorch wheel, Voicebox automatically selects the ROCm device. The _get_gpu_status function in backend/app.py handles detection automatically by checking torch.version.hip.

What if Voicebox detects my AMD GPU as CUDA?

This indicates PyTorch is not installed with ROCm support. Verify you installed PyTorch using the ROCm wheel index (https://download.pytorch.org/whl/rocm6.0) rather than the CUDA wheel index. The torch.version.hip attribute must return a version string for Voicebox to classify the device as ROCm.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →