How to Pin a Compute Device for TTS Generation (CUDA | MPS | ROCm | CPU)
Set the OMNIVOICE_TTS_DEVICE environment variable to cpu, cuda, cuda:<index>, mps, or rocm before launching VoiceStudio, and the backend loader binds all TTS inference to that exact accelerator.
VoiceStudio routes text-to-speech workloads through a device-aware backend that lets operators pin a compute device for TTS generation without touching source code. Whether you are running NVIDIA CUDA, Apple MPS, AMD ROCm, or CPU-only hosts, the selection logic is centralized in the engine loader and respects explicit overrides through standard environment variables.
How VoiceStudio Selects a TTS Compute Device
VoiceStudio’s engine loader applies a three-step resolution rule to decide where tensors are allocated and models are executed.
Step 1 — Override with OMNIVOICE_TTS_DEVICE
The loader first checks for an explicit override in the process environment. Setting OMNIVOICE_TTS_DEVICE — or OMNIVOICE_TTS_BACKEND for engine-level granularity — to a supported identifier forces the TTS backend to use that device regardless of automatic probing.
Supported identifiers include:
cpu— Force CPU execution.cuda— Use the default CUDA GPU.cuda:<index>— Pin to a specific GPU ordinal (for example,cuda:1).mps— Bind to the Apple Silicon MPS graph backend.rocm— Target an AMD ROCm accelerator.
Inside backend/services/tts_backend.py, the function runtime_compute_profile invokes detect_host_caps, where the variable is evaluated around lines 282–285 and translated into a torch.device object. If the variable is unset, the flow continues to automatic detection.
Step 2 — Automatic Accelerator Fallback
When no environment override exists, each engine probes the host runtime in a fixed priority order:
torch.cuda.is_available()returnsTrue→ selectcuda(orcuda:<index>when multiple GPUs are present).torch.backends.mps.is_available()returnsTrue→ selectmps.- ROCm detection is performed by
backend/engines/omnivoice_gguf/hardware_probe.py, which inspects the ROCm driver version and reports hardware familyrocm. - If none of the above succeed, the engine falls back to
torch.device("cpu").
This logic is coordinated through get_best_device() in backend/engines/omnivoice_subprocess/main.py at lines 145–151. That helper delegates to detect_host_caps() in core/device_caps.py to normalize the hardware family as cuda, mps, rocm, or cpu.
Step 3 — Mask GPUs with CUDA_VISIBLE_DEVICES
For multi-GPU workstations you can further restrict visibility at the process level. Setting CUDA_VISIBLE_DEVICES to a specific card index (for example, 2) hides all other GPUs from the runtime. When OMNIVOICE_TTS_DEVICE=cuda is combined with CUDA_VISIBLE_DEVICES=2, the loader resolves the default CUDA device to the first—and only—visible GPU, effectively pinning execution to that physical card.
Code Examples for Pinning TTS Devices
The following shell commands demonstrate how to pin a compute device for TTS generation before starting VoiceStudio:
# 1. Pin to the first CUDA GPU (index 0)
export OMNIVOICE_TTS_DEVICE=cuda:0
voice-studio
# 2. Pin to Apple Silicon MPS accelerator
export OMNIVOICE_TTS_DEVICE=mps
voice-studio
# 3. Pin to ROCm (AMD) accelerator
export OMNIVOICE_TTS_DEVICE=rocm
voice-studio
# 4. Force CPU execution (no GPU required)
export OMNIVOICE_TTS_DEVICE=cpu
voice-studio
# 5. Limit visible GPUs (e.g., use only GPU 2 on a 4-GPU machine)
export CUDA_VISIBLE_DEVICES=2
export OMNIVOICE_TTS_DEVICE=cuda # resolves to cuda:0 (the only visible GPU)
voice-studio
If you embed VoiceStudio as a library, set the variable in Python before importing the backend:
import os
# Choose GPU 1 explicitly
os.environ["OMNIVOICE_TTS_DEVICE"] = "cuda:1"
from backend.services import tts_backend
backend = tts_backend.get_active_tts_backend()
print(backend.runtime_compute_profile({}))
Core Source Files That Handle Device Selection
Understanding the source map helps when debugging why a specific device was chosen:
backend/services/tts_backend.py— Central TTS interface that readsOMNIVOICE_TTS_DEVICEand builds the runtime compute profile viaruntime_compute_profile.backend/engines/omnivoice_subprocess/main.py— Implementsget_best_device()around lines 145–151; it respects the env-var override and triggers automatic fallback when no override is set.backend/engines/omnivoice_gguf/hardware_probe.py— Detects ROCm capability by reading driver version metadata and reporting the hardware family asrocm.core/device_caps.py— Low-level host-capability detection shared across engines; normalizes device families for downstream routing.backend/engines/_asr_sidecar/main.py— Implements parallel device-override logic for the ASR sidecar and serves as a useful reference for how VoiceStudio handles compute pinning beyond TTS.
Summary
OMNIVOICE_TTS_DEVICEis the single environment variable that pins a compute device for TTS generation in VoiceStudio, acceptingcpu,cuda,cuda:<index>,mps, orrocm.- When the variable is absent, the engine auto-detects accelerators via
torch.cuda,torch.backends.mps, and the ROCm probe inhardware_probe.py, falling back to CPU if nothing is available. - Use
CUDA_VISIBLE_DEVICESto mask unwanted GPUs so that even a genericcudatarget resolves to a specific physical card. - All routing ultimately passes through
tts_backend.pyandomnivoice_subprocess/main.py, making device behavior transparent and reproducible across deployments.
Frequently Asked Questions
What environment variable pins the TTS device in VoiceStudio?
VoiceStudio uses OMNIVOICE_TTS_DEVICE to pin the TTS compute device. Set it before process startup to one of the supported identifiers, and backend/services/tts_backend.py will bind the backend to that device without further configuration.
Does VoiceStudio support multi-GPU CUDA indexing?
Yes. You can pass a specific ordinal such as cuda:1 or cuda:2 to OMNIVOICE_TTS_DEVICE. Alternatively, restrict GPU visibility with CUDA_VISIBLE_DEVICES and set the device variable to cuda; the loader will map cuda:0 to the first visible card.
Can I force CPU-only TTS generation even if a GPU is present?
Yes. Setting OMNIVOICE_TTS_DEVICE=cpu bypasses all accelerator detection and forces the engine to use torch.device("cpu"). This is useful for reproducible debugging or for running VoiceStudio on headless servers without display drivers.
How does VoiceStudio detect ROCm capability on AMD hardware?
The engine runs a dedicated probe inside backend/engines/omnivoice_gguf/hardware_probe.py that inspects ROCm driver version information. If the probe validates the ROCm runtime, it reports hardware family rocm, and the TTS loader targets the AMD accelerator accordingly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →