How to Configure VoiceStudio for Specific Hardware (CUDA, MPS, ROCm, CPU)

VoiceStudio automatically detects available compute hardware at startup and routes model inference to the most appropriate backend, with full support for manual overrides via CLI flags, environment variables, and persistent settings.

VoiceStudio is an open-source inference engine for voice synthesis. Understanding how to configure VoiceStudio for specific hardware is essential when deploying across diverse environments — from NVIDIA data center GPUs to Apple Silicon laptops and AMD ROCm workstations. This guide walks through the automatic detection system and every available override mechanism.

Automatic Hardware Detection

VoiceStudio probes the host system once at startup and caches the result for the process lifetime. The detection logic lives in three interconnected modules:

Module Responsibility Key Function
backend/core/device_caps.py Defines the immutable HostCaps dataclass and runs lazy probe detect_host_caps()
backend/engines/omnivoice_gguf/hardware_probe.py Builds HardwareCapabilities for the routing matrix detect_capabilities()
backend/services/model_manager.py Public entry point used by CLI, UI, and workers get_best_device()

Detection Priority Order

The probe executes checks in strict priority order:

  1. CUDA — torch.cuda.is_available() validates driver presence. Additional queries read torch.version.cuda and VRAM via torch.cuda.mem_get_info(). Result: backend="cuda", vram_gb populated.
  2. Apple MPS — torch.backends.mps.is_available() detects Apple Silicon. Memory is estimated as half the host's virtual memory. Result: backend="mps".
  3. ROCm — torch.version.hip and rocm-smi command presence confirm AMD driver installation. VRAM reads through the CUDA-compatible API. Result: backend="rocm".
  4. CPU fallback — When all GPU checks fail. Result: backend="cpu", vram_gb=0.

All detection occurs without network calls. The HostCaps dataclass caches the result, ensuring consistent device selection across all VoiceStudio components.

Override Methods for Manual Configuration

VoiceStudio provides five mechanisms to override automatic detection. These are validated — requesting an unavailable device triggers a warning and graceful fallback.

CLI Flag (Command-Scoped)

Force a specific device for a single command invocation:

voice-studio generate --text "Hello world" --device cuda

# Valid values: cuda, mps, rocm, cpu

The --device flag takes precedence over cached detection for that process only. Implemented in argument parsing before get_best_device() returns its value.

Settings UI (Persistent)

Store a default device preference across sessions:

  1. Navigate to Settings → Engines → Default device
  2. Select CUDA, MPS, ROCm, or CPU
  3. Save to ~/.voice_studio/config.json

get_best_device() reads this config file before falling back to hardware probing. The UI component stores the override at frontend/src/components/Settings/DeviceSelect.jsx.

Environment Variable (Process-Scoped)

Set VOICESTUDIO_DEVICE for script automation or containerized deployments:

export VOICESTUDIO_DEVICE=cpu
voice-studio generate --text "Hello world"

The model_manager.get_best_device() function checks this variable early in initialization. Unset it to restore automatic detection:

unset VOICESTUDIO_DEVICE

GPU Visibility (System-Scoped)

Limit GPU visibility without changing VoiceStudio configuration. Useful for multi-GPU hosts:


# Pin to CUDA device 1 only

export CUDA_VISIBLE_DEVICES=1
voice-studio generate --text "Hello world"

# Hide all GPUs, force CPU fallback

export CUDA_VISIBLE_DEVICES=""
voice-studio generate --text "Hello world"

ROCm equivalent:

export ROCM_VISIBLE_DEVICES=0

The probe respects these variables — detection runs against the filtered device list.

MPS Fallback Behavior

Enable automatic CPU fallback on macOS when MPS memory allocation fails:

export PYTORCH_ENABLE_MPS_FALLBACK=1

The backend still reports "mps", but PyTorch transparently moves failed allocations to CPU. This preserves the device string for consistency in logs and API responses.

Programmatic Device Inspection

Query or force devices from Python:


# Retrieve the currently selected device

from services.model_manager import get_best_device

print("Execution device:", get_best_device())

# Output: "cuda", "mps", "rocm", or "cpu"

Inspect full capabilities for debugging:

from backend.core.device_caps import detect_host_caps

caps = detect_host_caps()
print(caps)

# HostCaps(backend='cuda', vram_gb=23.7, gpu_name='NVIDIA RTX 3090', ...)

Force CPU for a subprocess without modifying global state:

import subprocess
import os

env = os.environ.copy()
env["VOICESTUDIO_DEVICE"] = "cpu"

subprocess.run(
    ["voice-studio", "generate", "--text", "Hello world"],
    env=env,
    check=True,
)

Pin a specific CUDA GPU on multi-GPU hosts:

import os

os.environ["CUDA_VISIBLE_DEVICES"] = "2"  # GPU index 2

from services.model_manager import get_best_device

device = get_best_device()  # Returns "cuda", using GPU 2

Hardware-Specific Configuration Scenarios

Environment Configuration
NVIDIA GPU (≥8 GB VRAM) No action required. Probe selects "cuda" automatically.
Apple Silicon (M1/M2/M3) No action required. Probe selects "mps". Force "cpu" on Intel Macs if MPS detection errors occur.
AMD GPU with ROCm Install ROCm driver and rocm-smi. Probe selects "rocm". Test CUDA fallback with CUDA_VISIBLE_DEVICES="".
Multi-GPU server Use CUDA_VISIBLE_DEVICES=N or ROCM_VISIBLE_DEVICES=N to pin to specific GPU.
Low-VRAM system (≤2 GB) Set --device cpu or VOICESTUDIO_DEVICE=cpu to prevent out-of-memory crashes.

Integration Points

The configured device value propagates through VoiceStudio's architecture:

Summary

  • Automatic detection probes CUDA → MPS → ROCm → CPU in sequence, caching results in HostCaps
  • Five override methods provide flexibility: CLI flags, UI settings, environment variables, GPU visibility masks, and MPS fallback flags
  • Validation and graceful fallback prevent hard failures when requesting unavailable hardware
  • Programmatic access via get_best_device() and detect_host_caps() enables custom integrations
  • Cross-platform consistency in backend strings ("cuda", "mps", "rocm", "cpu") simplifies deployment scripts

Frequently Asked Questions

What if VoiceStudio detects the wrong GPU?

Set CUDA_VISIBLE_DEVICES or ROCM_VISIBLE_DEVICES to limit GPU visibility before startup. For persistent fixes, use the Settings UI or VOICESTUDIO_DEVICE environment variable to force a specific backend.

Can I use VoiceStudio without any GPU?

Yes. The probe automatically falls back to "cpu" when no GPU drivers are detected. For explicit CPU-only operation, set --device cpu or VOICESTUDIO_DEVICE=cpu to bypass detection entirely.

Why does MPS show on my Intel Mac?

Apple's MPS backend may incorrectly report availability on some Intel Mac configurations. Force CPU mode with --device cpu if you encounter runtime errors or inconsistent behavior.

How do I verify which device VoiceStudio is actually using?

Check the startup logs for the HostCaps probe result, query get_best_device() programmatically, or call the /system REST endpoint. The backend string is consistent across all interfaces.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →