# How VoiceStudio Routes Inference Across CUDA, MPS, MLX, and ROCm: A Deep Dive into the GPU Gateway

> Discover how the VoiceStudio GPU gateway routes inference across CUDA, MPS, MLX, and ROCm by detecting capabilities and prioritizing backends. Learn about local execution and remote dispatch.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: deep-dive
- Published: 2026-09-13

---

**The VoiceStudio GPU gateway routes inference requests by detecting host capabilities, selecting compatible backends based on a strict priority order, and deciding between local execution or remote dispatch using the `decide()` and `run()` methods in [`services/gpu_gateway.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/gpu_gateway.py).**

VoiceStudio abstracts hardware complexity behind a unified **GPU gateway** that intelligently handles text-to-speech (TTS) and automatic speech recognition (ASR) workloads across heterogeneous compute environments. This routing system examines the host's available accelerators—CUDA, ROCm, Apple MPS, or MLX—and matches them against engine compatibility lists to determine optimal execution paths. Understanding this architecture helps developers optimize inference latency and resource utilization in both on-premise and distributed deployments.

## Step-by-Step Routing Logic

The gateway follows a deterministic pipeline when processing inference requests, from hardware detection to final execution.

### 1. Host Capability Detection

The routing process begins in [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py), where the system probes runtime capabilities using `torch.cuda.is_available()` and platform-specific flags. This module constructs a `HostCaps` object that catalogs available accelerators and estimates VRAM budgets.

The detection follows a strict priority hierarchy: **CUDA/ROCm > Intel XPU > DirectML > MPS > CPU**. For ROCm systems, the code specifically examines GPU architecture strings (e.g., `gfx...`) to distinguish genuine ROCm support from CUDA-compatible builds that masquerade as CUDA devices.

```python

# Simplified capability detection flow

from backend.services.model_manager import ModelManager

manager = ModelManager()
caps = manager.get_host_caps()  # Returns HostCaps object

print(caps.vram_gb)             # Available GPU memory or system RAM/2 for MPS

print(caps.cuda_available)      # True for both CUDA and ROCm

print(caps.mlx_supported)       # True only on Apple Silicon with Metal

```

### 2. Backend Class Selection

Once capabilities are established, [`backend/services/engine_routing.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py) maps the engine's declared compatibility list to concrete backend implementations. The router selects from classes such as `OmniVoiceBackend` for CUDA, `OmniVoiceMPSSubprocessBackend` for Apple Silicon, or `MLXAudioBackend` for MLX-capable systems.

If the required backend cannot satisfy the request—for instance, when a CUDA-only engine encounters an MPS-only host—the system flags a fallback condition before proceeding to the decision phase.

### 3. Local vs. Remote Decision

The core decision logic resides in [`services/gpu_gateway.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/gpu_gateway.py) within the `decide()` method. This function returns one of three states: `LOCAL_CHOSEN`, `LOCAL_FALLBACK`, or `REMOTE`, based on whether the selected backend can execute on the current host.

```python
from services import gpu_gateway

# Evaluate execution strategy for a TTS request

decision = gpu_gateway.decide("tts")  

# Returns: Decision.LOCAL_CHOSEN, Decision.LOCAL_FALLBACK, or Decision.REMOTE

```

If the host lacks the necessary hardware (e.g., requesting CUDA inference on an MPS-only MacBook), the gateway returns `LOCAL_FALLBACK` to trigger CPU execution or `REMOTE` to dispatch the job to a GPU worker cluster via Docker or Tailscale nodes.

### 4. Job Execution and Staging

The `run()` method in [`gpu_gateway.py`](https://github.com/debpalash/VoiceStudio/blob/main/gpu_gateway.py) orchestrates final execution by creating a `JobRun` object that stages input artifacts. Depending on the decision, it invokes either `_run_local()` for direct backend invocation or `_run_remote()` to transmit the payload to a control-plane GPU worker.

```python

# Construct local and remote execution paths

local_call = gpu_gateway.LocalCall(
    fn=lambda: tts_backend.generate(text, voice),
    what="TTS generate"
)

remote_call = gpu_gateway.RemoteCall(
    fn="tts/generate",
    payload={"text": text, "voice": voice}
)

# Execute based on previous decision

result = await gpu_gateway.run(
    local=local_call, 
    remote=remote_call, 
    decision=decision
)

```

## Platform-Specific Routing Behaviors

Different GPU stacks require unique handling patterns within the gateway architecture.

### CUDA and ROCm Handling

For NVIDIA and AMD hardware, the gateway treats `torch.cuda.is_available()` as a preliminary check. On ROCm systems, the code validates device architecture lists to confirm ROCm compatibility. If the architecture is unsupported, the gateway treats the host as CUDA-compatible but without ROCm acceleration, falling back to CPU execution or CUDA-only backends as appropriate.

### Apple Silicon MPS Routing

MPS (Metal Performance Shaders) operates as a unified-memory backend on Apple Silicon. The gateway creates a **single-worker pool** with `HostCaps.vram_gb` set to half the system RAM. All MPS jobs route through `OmniVoiceMPSSubprocessBackend`, and the gateway disables self-deadlocking patterns by ensuring the single-worker constraint prevents resource contention.

### MLX Backend Selection

MLX acceleration is strictly limited to Apple Silicon devices with Metal support. The `core.device_caps.mlx_supported()` function returns `false` on non-Apple hardware, preventing gateway selection of `MLXAudioBackend` or `MLXWhisperBackend` on incompatible systems. When available, MLX backends execute locally without remote fallback options, leveraging the unified memory architecture for zero-copy tensor operations.

## Complete End-to-End Implementation

The following pattern demonstrates production usage of the VoiceStudio GPU gateway for cross-platform inference:

```python
import asyncio
from services import gpu_gateway
from backend.services import tts_backend

async def generate_speech(text: str, voice: str):
    # Step 1: Gateway decides execution path based on hardware

    decision = gpu_gateway.decide("tts")
    
    # Step 2: Prepare local execution wrapper

    local = gpu_gateway.LocalCall(
        fn=lambda: tts_backend.generate(text, voice),
        what="TTS generation"
    )
    
    # Step 3: Prepare remote fallback (GPU worker cluster)

    remote = gpu_gateway.RemoteCall(
        fn="tts/generate",
        payload={"text": text, "voice": voice, "priority": "high"}
    )
    
    # Step 4: Execute via gateway

    result = await gpu_gateway.run(
        local=local,
        remote=remote,
        decision=decision
    )
    
    # Step 5: Decode audio artifact

    waveform, sample_rate = gpu_gateway.decode_audio_artifact(result)
    return waveform, sample_rate

# Runtime behavior:

# - CUDA/ROCm host: Runs locally on GPU

# - MPS-only Mac: Routes to OmniVoiceMPSSubprocessBackend

# - MLX-capable Mac: Uses MLXAudioBackend

# - No GPU: REMOTE dispatch to GPU worker or CPU fallback

```

## Summary

- **Capability Detection**: The system probes hardware using [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py) with a priority order of CUDA/ROCm > Intel XPU > DirectML > MPS > CPU.
- **Routing Decision**: The `decide()` method in [`services/gpu_gateway.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/gpu_gateway.py) selects between `LOCAL_CHOSEN`, `LOCAL_FALLBACK`, and `REMOTE` execution strategies.
- **Backend Abstraction**: Concrete classes like `OmniVoiceBackend`, `OmniVoiceMPSSubprocessBackend`, and `MLXAudioBackend` handle platform-specific implementations.
- **ROCm Validation**: The gateway verifies `gfx` architectures to distinguish true ROCm support from CUDA-compatible builds.
- **MPS Constraints**: Apple Silicon routing uses a single-worker pool with VRAM budgets derived from system RAM to prevent deadlocks.
- **MLX Availability**: MLX backends are exclusively selected on Apple Silicon with Metal support, as enforced by `core.device_caps.mlx_supported()`.

## Frequently Asked Questions

### How does VoiceStudio handle ROCm devices that report as CUDA-compatible?

VoiceStudio detects ROCm hardware by checking GPU architecture strings (e.g., `gfx...`) in [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py). While ROCm builds report `torch.cuda.is_available() == True`, the gateway validates the specific architecture against supported ROCm configurations. If validation fails, the system treats the device as CUDA-compatible but without ROCm acceleration, falling back to CPU execution or standard CUDA backends.

### What happens when a CUDA-only engine runs on an Apple Silicon Mac?

When `gpu_gateway.decide()` detects a CUDA-only engine on an MPS-only host, it returns `LOCAL_FALLBACK` or `REMOTE` depending on configuration. The `LOCAL_FALLBACK` state triggers CPU execution, while `REMOTE` dispatches the request to a GPU worker service running on CUDA-capable hardware via the control-plane scheduler.

### Why does MPS routing use a single-worker pool on Apple Silicon?

The MPS backend creates a single-worker pool with `HostCaps.vram_gb` set to half the system RAM to prevent self-deadlocking patterns. Because MPS shares unified memory between CPU and GPU, concurrent workers could exhaust memory or create circular dependencies. The `OmniVoiceMPSSubprocessBackend` enforces this constraint to ensure stable inference on macOS devices.

### Can MLX backends fall back to remote execution?

No. When `core.device_caps.mlx_supported()` returns true on Apple Silicon, the gateway selects `MLXAudioBackend` or `MLXWhisperBackend` for local execution only. MLX does not support remote dispatch because the framework is exclusive to Apple Silicon architecture. If MLX initialization fails, the gateway falls back to MPS or CPU rather than remote GPU workers.