# How to Configure VoiceStudio for Specific Hardware (CUDA, MPS, ROCm, CPU)

> Configure VoiceStudio for CUDA, MPS, ROCm, or CPU. Easily manage hardware acceleration for your models with CLI flags and environment variables. Optimize performance now.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-10

---

**VoiceStudio automatically detects available compute hardware at startup and routes model inference to the most appropriate backend, with full support for manual overrides via CLI flags, environment variables, and persistent settings.**

VoiceStudio is an open-source inference engine for voice synthesis. Understanding how to configure VoiceStudio for specific hardware is essential when deploying across diverse environments — from NVIDIA data center GPUs to Apple Silicon laptops and AMD ROCm workstations. This guide walks through the automatic detection system and every available override mechanism.

## Automatic Hardware Detection

VoiceStudio probes the host system once at startup and caches the result for the process lifetime. The detection logic lives in three interconnected modules:

| Module | Responsibility | Key Function |
|--------|----------------|--------------|
| [`backend/core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/device_caps.py) | Defines the immutable `HostCaps` dataclass and runs lazy probe | `detect_host_caps()` |
| [`backend/engines/omnivoice_gguf/hardware_probe.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_gguf/hardware_probe.py) | Builds `HardwareCapabilities` for the routing matrix | `detect_capabilities()` |
| [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py) | Public entry point used by CLI, UI, and workers | `get_best_device()` |

### Detection Priority Order

The probe executes checks in strict priority order:

1. **CUDA** — `torch.cuda.is_available()` validates driver presence. Additional queries read `torch.version.cuda` and VRAM via `torch.cuda.mem_get_info()`. Result: `backend="cuda"`, `vram_gb` populated.
2. **Apple MPS** — `torch.backends.mps.is_available()` detects Apple Silicon. Memory is estimated as half the host's virtual memory. Result: `backend="mps"`.
3. **ROCm** — `torch.version.hip` and `rocm-smi` command presence confirm AMD driver installation. VRAM reads through the CUDA-compatible API. Result: `backend="rocm"`.
4. **CPU fallback** — When all GPU checks fail. Result: `backend="cpu"`, `vram_gb=0`.

All detection occurs **without network calls**. The `HostCaps` dataclass caches the result, ensuring consistent device selection across all VoiceStudio components.

## Override Methods for Manual Configuration

VoiceStudio provides five mechanisms to override automatic detection. These are validated — requesting an unavailable device triggers a warning and graceful fallback.

### CLI Flag (Command-Scoped)

Force a specific device for a single command invocation:

```bash
voice-studio generate --text "Hello world" --device cuda

# Valid values: cuda, mps, rocm, cpu

```

The `--device` flag takes precedence over cached detection for that process only. Implemented in argument parsing before `get_best_device()` returns its value.

### Settings UI (Persistent)

Store a default device preference across sessions:

1. Navigate to **Settings → Engines → Default device**
2. Select **CUDA**, **MPS**, **ROCm**, or **CPU**
3. Save to `~/.voice_studio/config.json`

`get_best_device()` reads this config file before falling back to hardware probing. The UI component stores the override at [`frontend/src/components/Settings/DeviceSelect.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/frontend/src/components/Settings/DeviceSelect.jsx).

### Environment Variable (Process-Scoped)

Set `VOICESTUDIO_DEVICE` for script automation or containerized deployments:

```bash
export VOICESTUDIO_DEVICE=cpu
voice-studio generate --text "Hello world"

```

The `model_manager.get_best_device()` function checks this variable early in initialization. Unset it to restore automatic detection:

```bash
unset VOICESTUDIO_DEVICE

```

### GPU Visibility (System-Scoped)

Limit GPU visibility without changing VoiceStudio configuration. Useful for multi-GPU hosts:

```bash

# Pin to CUDA device 1 only

export CUDA_VISIBLE_DEVICES=1
voice-studio generate --text "Hello world"

# Hide all GPUs, force CPU fallback

export CUDA_VISIBLE_DEVICES=""
voice-studio generate --text "Hello world"

```

ROCm equivalent:

```bash
export ROCM_VISIBLE_DEVICES=0

```

The probe respects these variables — detection runs against the filtered device list.

### MPS Fallback Behavior

Enable automatic CPU fallback on macOS when MPS memory allocation fails:

```bash
export PYTORCH_ENABLE_MPS_FALLBACK=1

```

The backend still reports `"mps"`, but PyTorch transparently moves failed allocations to CPU. This preserves the device string for consistency in logs and API responses.

## Programmatic Device Inspection

Query or force devices from Python:

```python

# Retrieve the currently selected device

from services.model_manager import get_best_device

print("Execution device:", get_best_device())

# Output: "cuda", "mps", "rocm", or "cpu"

```

Inspect full capabilities for debugging:

```python
from backend.core.device_caps import detect_host_caps

caps = detect_host_caps()
print(caps)

# HostCaps(backend='cuda', vram_gb=23.7, gpu_name='NVIDIA RTX 3090', ...)

```

Force CPU for a subprocess without modifying global state:

```python
import subprocess
import os

env = os.environ.copy()
env["VOICESTUDIO_DEVICE"] = "cpu"

subprocess.run(
    ["voice-studio", "generate", "--text", "Hello world"],
    env=env,
    check=True,
)

```

Pin a specific CUDA GPU on multi-GPU hosts:

```python
import os

os.environ["CUDA_VISIBLE_DEVICES"] = "2"  # GPU index 2

from services.model_manager import get_best_device

device = get_best_device()  # Returns "cuda", using GPU 2

```

## Hardware-Specific Configuration Scenarios

| Environment | Configuration |
|-------------|---------------|
| **NVIDIA GPU (≥8 GB VRAM)** | No action required. Probe selects `"cuda"` automatically. |
| **Apple Silicon (M1/M2/M3)** | No action required. Probe selects `"mps"`. Force `"cpu"` on Intel Macs if MPS detection errors occur. |
| **AMD GPU with ROCm** | Install ROCm driver and `rocm-smi`. Probe selects `"rocm"`. Test CUDA fallback with `CUDA_VISIBLE_DEVICES=""`. |
| **Multi-GPU server** | Use `CUDA_VISIBLE_DEVICES=N` or `ROCM_VISIBLE_DEVICES=N` to pin to specific GPU. |
| **Low-VRAM system (≤2 GB)** | Set `--device cpu` or `VOICESTUDIO_DEVICE=cpu` to prevent out-of-memory crashes. |

## Integration Points

The configured device value propagates through VoiceStudio's architecture:

- **REST API** — `/system` endpoint (defined in [`backend/api/routers/system.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/system.py)) exposes `backend` to monitoring tools
- **Routing matrix** — [`docs/specs/21-gpu-compat-matrix.md`](https://github.com/debpalash/VoiceStudio/blob/main/docs/specs/21-gpu-compat-matrix.md) maps `HostCaps.backend` to supported engine quantizations
- **Worker scheduler** — Distributes inference jobs based on the cached capability object

## Summary

- **Automatic detection** probes CUDA → MPS → ROCm → CPU in sequence, caching results in `HostCaps`
- **Five override methods** provide flexibility: CLI flags, UI settings, environment variables, GPU visibility masks, and MPS fallback flags
- **Validation and graceful fallback** prevent hard failures when requesting unavailable hardware
- **Programmatic access** via `get_best_device()` and `detect_host_caps()` enables custom integrations
- **Cross-platform consistency** in backend strings (`"cuda"`, `"mps"`, `"rocm"`, `"cpu"`) simplifies deployment scripts

## Frequently Asked Questions

### What if VoiceStudio detects the wrong GPU?

Set `CUDA_VISIBLE_DEVICES` or `ROCM_VISIBLE_DEVICES` to limit GPU visibility before startup. For persistent fixes, use the Settings UI or `VOICESTUDIO_DEVICE` environment variable to force a specific backend.

### Can I use VoiceStudio without any GPU?

Yes. The probe automatically falls back to `"cpu"` when no GPU drivers are detected. For explicit CPU-only operation, set `--device cpu` or `VOICESTUDIO_DEVICE=cpu` to bypass detection entirely.

### Why does MPS show on my Intel Mac?

Apple's MPS backend may incorrectly report availability on some Intel Mac configurations. Force CPU mode with `--device cpu` if you encounter runtime errors or inconsistent behavior.

### How do I verify which device VoiceStudio is actually using?

Check the startup logs for the `HostCaps` probe result, query `get_best_device()` programmatically, or call the `/system` REST endpoint. The backend string is consistent across all interfaces.