How to Pin a Compute Device for TTS Generation (CUDA | MPS | ROCm | CPU)

Set the OMNIVOICE_TTS_DEVICE environment variable to cpu, cuda, cuda:<index>, mps, or rocm before launching VoiceStudio, and the backend loader binds all TTS inference to that exact accelerator.

VoiceStudio routes text-to-speech workloads through a device-aware backend that lets operators pin a compute device for TTS generation without touching source code. Whether you are running NVIDIA CUDA, Apple MPS, AMD ROCm, or CPU-only hosts, the selection logic is centralized in the engine loader and respects explicit overrides through standard environment variables.

How VoiceStudio Selects a TTS Compute Device

VoiceStudio’s engine loader applies a three-step resolution rule to decide where tensors are allocated and models are executed.

Step 1 — Override with OMNIVOICE_TTS_DEVICE

The loader first checks for an explicit override in the process environment. Setting OMNIVOICE_TTS_DEVICE — or OMNIVOICE_TTS_BACKEND for engine-level granularity — to a supported identifier forces the TTS backend to use that device regardless of automatic probing.

Supported identifiers include:

  • cpu — Force CPU execution.
  • cuda — Use the default CUDA GPU.
  • cuda:<index> — Pin to a specific GPU ordinal (for example, cuda:1).
  • mps — Bind to the Apple Silicon MPS graph backend.
  • rocm — Target an AMD ROCm accelerator.

Inside backend/services/tts_backend.py, the function runtime_compute_profile invokes detect_host_caps, where the variable is evaluated around lines 282–285 and translated into a torch.device object. If the variable is unset, the flow continues to automatic detection.

Step 2 — Automatic Accelerator Fallback

When no environment override exists, each engine probes the host runtime in a fixed priority order:

  1. torch.cuda.is_available() returns True → select cuda (or cuda:<index> when multiple GPUs are present).
  2. torch.backends.mps.is_available() returns True → select mps.
  3. ROCm detection is performed by backend/engines/omnivoice_gguf/hardware_probe.py, which inspects the ROCm driver version and reports hardware family rocm.
  4. If none of the above succeed, the engine falls back to torch.device("cpu").

This logic is coordinated through get_best_device() in backend/engines/omnivoice_subprocess/main.py at lines 145–151. That helper delegates to detect_host_caps() in core/device_caps.py to normalize the hardware family as cuda, mps, rocm, or cpu.

Step 3 — Mask GPUs with CUDA_VISIBLE_DEVICES

For multi-GPU workstations you can further restrict visibility at the process level. Setting CUDA_VISIBLE_DEVICES to a specific card index (for example, 2) hides all other GPUs from the runtime. When OMNIVOICE_TTS_DEVICE=cuda is combined with CUDA_VISIBLE_DEVICES=2, the loader resolves the default CUDA device to the first—and only—visible GPU, effectively pinning execution to that physical card.

Code Examples for Pinning TTS Devices

The following shell commands demonstrate how to pin a compute device for TTS generation before starting VoiceStudio:


# 1. Pin to the first CUDA GPU (index 0)

export OMNIVOICE_TTS_DEVICE=cuda:0
voice-studio

# 2. Pin to Apple Silicon MPS accelerator

export OMNIVOICE_TTS_DEVICE=mps
voice-studio

# 3. Pin to ROCm (AMD) accelerator

export OMNIVOICE_TTS_DEVICE=rocm
voice-studio

# 4. Force CPU execution (no GPU required)

export OMNIVOICE_TTS_DEVICE=cpu
voice-studio

# 5. Limit visible GPUs (e.g., use only GPU 2 on a 4-GPU machine)

export CUDA_VISIBLE_DEVICES=2
export OMNIVOICE_TTS_DEVICE=cuda   # resolves to cuda:0 (the only visible GPU)

voice-studio

If you embed VoiceStudio as a library, set the variable in Python before importing the backend:

import os

# Choose GPU 1 explicitly

os.environ["OMNIVOICE_TTS_DEVICE"] = "cuda:1"

from backend.services import tts_backend

backend = tts_backend.get_active_tts_backend()
print(backend.runtime_compute_profile({}))

Core Source Files That Handle Device Selection

Understanding the source map helps when debugging why a specific device was chosen:

Summary

  • OMNIVOICE_TTS_DEVICE is the single environment variable that pins a compute device for TTS generation in VoiceStudio, accepting cpu, cuda, cuda:<index>, mps, or rocm.
  • When the variable is absent, the engine auto-detects accelerators via torch.cuda, torch.backends.mps, and the ROCm probe in hardware_probe.py, falling back to CPU if nothing is available.
  • Use CUDA_VISIBLE_DEVICES to mask unwanted GPUs so that even a generic cuda target resolves to a specific physical card.
  • All routing ultimately passes through tts_backend.py and omnivoice_subprocess/main.py, making device behavior transparent and reproducible across deployments.

Frequently Asked Questions

What environment variable pins the TTS device in VoiceStudio?

VoiceStudio uses OMNIVOICE_TTS_DEVICE to pin the TTS compute device. Set it before process startup to one of the supported identifiers, and backend/services/tts_backend.py will bind the backend to that device without further configuration.

Does VoiceStudio support multi-GPU CUDA indexing?

Yes. You can pass a specific ordinal such as cuda:1 or cuda:2 to OMNIVOICE_TTS_DEVICE. Alternatively, restrict GPU visibility with CUDA_VISIBLE_DEVICES and set the device variable to cuda; the loader will map cuda:0 to the first visible card.

Can I force CPU-only TTS generation even if a GPU is present?

Yes. Setting OMNIVOICE_TTS_DEVICE=cpu bypasses all accelerator detection and forces the engine to use torch.device("cpu"). This is useful for reproducible debugging or for running VoiceStudio on headless servers without display drivers.

How does VoiceStudio detect ROCm capability on AMD hardware?

The engine runs a dedicated probe inside backend/engines/omnivoice_gguf/hardware_probe.py that inspects ROCm driver version information. If the probe validates the ROCm runtime, it reports hardware family rocm, and the TTS loader targets the AMD accelerator accordingly.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →