# How to Pin a Compute Device for TTS Generation (CUDA | MPS | ROCm | CPU)

> Easily pin your compute device CPU CUDA MPS ROCm for TTS generation in VoiceStudio by setting the OMNIVOICE_TTS_DEVICE environment variable. Optimize your inference now.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: how-to-guide
- Published: 2026-09-09

---

**Set the `OMNIVOICE_TTS_DEVICE` environment variable to `cpu`, `cuda`, `cuda:<index>`, `mps`, or `rocm` before launching VoiceStudio, and the backend loader binds all TTS inference to that exact accelerator.**

VoiceStudio routes text-to-speech workloads through a device-aware backend that lets operators pin a compute device for TTS generation without touching source code. Whether you are running NVIDIA CUDA, Apple MPS, AMD ROCm, or CPU-only hosts, the selection logic is centralized in the engine loader and respects explicit overrides through standard environment variables.

## How VoiceStudio Selects a TTS Compute Device

VoiceStudio’s engine loader applies a three-step resolution rule to decide where tensors are allocated and models are executed.

### Step 1 — Override with `OMNIVOICE_TTS_DEVICE`

The loader first checks for an explicit override in the process environment. Setting **`OMNIVOICE_TTS_DEVICE`** — or **`OMNIVOICE_TTS_BACKEND`** for engine-level granularity — to a supported identifier forces the TTS backend to use that device regardless of automatic probing.

Supported identifiers include:

- `cpu` — Force CPU execution.
- `cuda` — Use the default CUDA GPU.
- `cuda:<index>` — Pin to a specific GPU ordinal (for example, `cuda:1`).
- `mps` — Bind to the Apple Silicon MPS graph backend.
- `rocm` — Target an AMD ROCm accelerator.

Inside [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py), the function `runtime_compute_profile` invokes `detect_host_caps`, where the variable is evaluated around lines 282–285 and translated into a `torch.device` object. If the variable is unset, the flow continues to automatic detection.

### Step 2 — Automatic Accelerator Fallback

When no environment override exists, each engine probes the host runtime in a fixed priority order:

1. **`torch.cuda.is_available()`** returns `True` → select `cuda` (or `cuda:<index>` when multiple GPUs are present).
2. **`torch.backends.mps.is_available()`** returns `True` → select `mps`.
3. **ROCm detection** is performed by [`backend/engines/omnivoice_gguf/hardware_probe.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_gguf/hardware_probe.py), which inspects the ROCm driver version and reports hardware family `rocm`.
4. If none of the above succeed, the engine falls back to `torch.device("cpu")`.

This logic is coordinated through `get_best_device()` in [`backend/engines/omnivoice_subprocess/main.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_subprocess/main.py) at lines 145–151. That helper delegates to `detect_host_caps()` in [`core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/device_caps.py) to normalize the hardware family as `cuda`, `mps`, `rocm`, or `cpu`.

### Step 3 — Mask GPUs with `CUDA_VISIBLE_DEVICES`

For multi-GPU workstations you can further restrict visibility at the process level. Setting `CUDA_VISIBLE_DEVICES` to a specific card index (for example, `2`) hides all other GPUs from the runtime. When `OMNIVOICE_TTS_DEVICE=cuda` is combined with `CUDA_VISIBLE_DEVICES=2`, the loader resolves the default CUDA device to the first—and only—visible GPU, effectively pinning execution to that physical card.

## Code Examples for Pinning TTS Devices

The following shell commands demonstrate how to pin a compute device for TTS generation before starting VoiceStudio:

```bash

# 1. Pin to the first CUDA GPU (index 0)

export OMNIVOICE_TTS_DEVICE=cuda:0
voice-studio

```

```bash

# 2. Pin to Apple Silicon MPS accelerator

export OMNIVOICE_TTS_DEVICE=mps
voice-studio

```

```bash

# 3. Pin to ROCm (AMD) accelerator

export OMNIVOICE_TTS_DEVICE=rocm
voice-studio

```

```bash

# 4. Force CPU execution (no GPU required)

export OMNIVOICE_TTS_DEVICE=cpu
voice-studio

```

```bash

# 5. Limit visible GPUs (e.g., use only GPU 2 on a 4-GPU machine)

export CUDA_VISIBLE_DEVICES=2
export OMNIVOICE_TTS_DEVICE=cuda   # resolves to cuda:0 (the only visible GPU)

voice-studio

```

If you embed VoiceStudio as a library, set the variable in Python before importing the backend:

```python
import os

# Choose GPU 1 explicitly

os.environ["OMNIVOICE_TTS_DEVICE"] = "cuda:1"

from backend.services import tts_backend

backend = tts_backend.get_active_tts_backend()
print(backend.runtime_compute_profile({}))

```

## Core Source Files That Handle Device Selection

Understanding the source map helps when debugging why a specific device was chosen:

- **[`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py)** — Central TTS interface that reads `OMNIVOICE_TTS_DEVICE` and builds the runtime compute profile via `runtime_compute_profile`.
- **[`backend/engines/omnivoice_subprocess/main.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_subprocess/main.py)** — Implements `get_best_device()` around lines 145–151; it respects the env-var override and triggers automatic fallback when no override is set.
- **[`backend/engines/omnivoice_gguf/hardware_probe.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_gguf/hardware_probe.py)** — Detects ROCm capability by reading driver version metadata and reporting the hardware family as `rocm`.
- **[`core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/device_caps.py)** — Low-level host-capability detection shared across engines; normalizes device families for downstream routing.
- **[`backend/engines/_asr_sidecar/main.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/_asr_sidecar/main.py)** — Implements parallel device-override logic for the ASR sidecar and serves as a useful reference for how VoiceStudio handles compute pinning beyond TTS.

## Summary

- **`OMNIVOICE_TTS_DEVICE`** is the single environment variable that pins a compute device for TTS generation in VoiceStudio, accepting `cpu`, `cuda`, `cuda:<index>`, `mps`, or `rocm`.
- When the variable is absent, the engine auto-detects accelerators via `torch.cuda`, `torch.backends.mps`, and the ROCm probe in [`hardware_probe.py`](https://github.com/debpalash/VoiceStudio/blob/main/hardware_probe.py), falling back to CPU if nothing is available.
- Use **`CUDA_VISIBLE_DEVICES`** to mask unwanted GPUs so that even a generic `cuda` target resolves to a specific physical card.
- All routing ultimately passes through [`tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/tts_backend.py) and [`omnivoice_subprocess/main.py`](https://github.com/debpalash/VoiceStudio/blob/main/omnivoice_subprocess/main.py), making device behavior transparent and reproducible across deployments.

## Frequently Asked Questions

### What environment variable pins the TTS device in VoiceStudio?

VoiceStudio uses **`OMNIVOICE_TTS_DEVICE`** to pin the TTS compute device. Set it before process startup to one of the supported identifiers, and [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py) will bind the backend to that device without further configuration.

### Does VoiceStudio support multi-GPU CUDA indexing?

Yes. You can pass a specific ordinal such as `cuda:1` or `cuda:2` to `OMNIVOICE_TTS_DEVICE`. Alternatively, restrict GPU visibility with `CUDA_VISIBLE_DEVICES` and set the device variable to `cuda`; the loader will map `cuda:0` to the first visible card.

### Can I force CPU-only TTS generation even if a GPU is present?

Yes. Setting `OMNIVOICE_TTS_DEVICE=cpu` bypasses all accelerator detection and forces the engine to use `torch.device("cpu")`. This is useful for reproducible debugging or for running VoiceStudio on headless servers without display drivers.

### How does VoiceStudio detect ROCm capability on AMD hardware?

The engine runs a dedicated probe inside **[`backend/engines/omnivoice_gguf/hardware_probe.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_gguf/hardware_probe.py)** that inspects ROCm driver version information. If the probe validates the ROCm runtime, it reports hardware family `rocm`, and the TTS loader targets the AMD accelerator accordingly.