How VoiceStudio Routes Inference Across CUDA, MPS, MLX, and ROCm: A Deep Dive into the GPU Gateway

The VoiceStudio GPU gateway routes inference requests by detecting host capabilities, selecting compatible backends based on a strict priority order, and deciding between local execution or remote dispatch using the decide() and run() methods in services/gpu_gateway.py.

VoiceStudio abstracts hardware complexity behind a unified GPU gateway that intelligently handles text-to-speech (TTS) and automatic speech recognition (ASR) workloads across heterogeneous compute environments. This routing system examines the host's available accelerators—CUDA, ROCm, Apple MPS, or MLX—and matches them against engine compatibility lists to determine optimal execution paths. Understanding this architecture helps developers optimize inference latency and resource utilization in both on-premise and distributed deployments.

Step-by-Step Routing Logic

The gateway follows a deterministic pipeline when processing inference requests, from hardware detection to final execution.

1. Host Capability Detection

The routing process begins in backend/services/model_manager.py, where the system probes runtime capabilities using torch.cuda.is_available() and platform-specific flags. This module constructs a HostCaps object that catalogs available accelerators and estimates VRAM budgets.

The detection follows a strict priority hierarchy: CUDA/ROCm > Intel XPU > DirectML > MPS > CPU. For ROCm systems, the code specifically examines GPU architecture strings (e.g., gfx...) to distinguish genuine ROCm support from CUDA-compatible builds that masquerade as CUDA devices.


# Simplified capability detection flow

from backend.services.model_manager import ModelManager

manager = ModelManager()
caps = manager.get_host_caps()  # Returns HostCaps object

print(caps.vram_gb)             # Available GPU memory or system RAM/2 for MPS

print(caps.cuda_available)      # True for both CUDA and ROCm

print(caps.mlx_supported)       # True only on Apple Silicon with Metal

2. Backend Class Selection

Once capabilities are established, backend/services/engine_routing.py maps the engine's declared compatibility list to concrete backend implementations. The router selects from classes such as OmniVoiceBackend for CUDA, OmniVoiceMPSSubprocessBackend for Apple Silicon, or MLXAudioBackend for MLX-capable systems.

If the required backend cannot satisfy the request—for instance, when a CUDA-only engine encounters an MPS-only host—the system flags a fallback condition before proceeding to the decision phase.

3. Local vs. Remote Decision

The core decision logic resides in services/gpu_gateway.py within the decide() method. This function returns one of three states: LOCAL_CHOSEN, LOCAL_FALLBACK, or REMOTE, based on whether the selected backend can execute on the current host.

from services import gpu_gateway

# Evaluate execution strategy for a TTS request

decision = gpu_gateway.decide("tts")  

# Returns: Decision.LOCAL_CHOSEN, Decision.LOCAL_FALLBACK, or Decision.REMOTE

If the host lacks the necessary hardware (e.g., requesting CUDA inference on an MPS-only MacBook), the gateway returns LOCAL_FALLBACK to trigger CPU execution or REMOTE to dispatch the job to a GPU worker cluster via Docker or Tailscale nodes.

4. Job Execution and Staging

The run() method in gpu_gateway.py orchestrates final execution by creating a JobRun object that stages input artifacts. Depending on the decision, it invokes either _run_local() for direct backend invocation or _run_remote() to transmit the payload to a control-plane GPU worker.


# Construct local and remote execution paths

local_call = gpu_gateway.LocalCall(
    fn=lambda: tts_backend.generate(text, voice),
    what="TTS generate"
)

remote_call = gpu_gateway.RemoteCall(
    fn="tts/generate",
    payload={"text": text, "voice": voice}
)

# Execute based on previous decision

result = await gpu_gateway.run(
    local=local_call, 
    remote=remote_call, 
    decision=decision
)

Platform-Specific Routing Behaviors

Different GPU stacks require unique handling patterns within the gateway architecture.

CUDA and ROCm Handling

For NVIDIA and AMD hardware, the gateway treats torch.cuda.is_available() as a preliminary check. On ROCm systems, the code validates device architecture lists to confirm ROCm compatibility. If the architecture is unsupported, the gateway treats the host as CUDA-compatible but without ROCm acceleration, falling back to CPU execution or CUDA-only backends as appropriate.

Apple Silicon MPS Routing

MPS (Metal Performance Shaders) operates as a unified-memory backend on Apple Silicon. The gateway creates a single-worker pool with HostCaps.vram_gb set to half the system RAM. All MPS jobs route through OmniVoiceMPSSubprocessBackend, and the gateway disables self-deadlocking patterns by ensuring the single-worker constraint prevents resource contention.

MLX Backend Selection

MLX acceleration is strictly limited to Apple Silicon devices with Metal support. The core.device_caps.mlx_supported() function returns false on non-Apple hardware, preventing gateway selection of MLXAudioBackend or MLXWhisperBackend on incompatible systems. When available, MLX backends execute locally without remote fallback options, leveraging the unified memory architecture for zero-copy tensor operations.

Complete End-to-End Implementation

The following pattern demonstrates production usage of the VoiceStudio GPU gateway for cross-platform inference:

import asyncio
from services import gpu_gateway
from backend.services import tts_backend

async def generate_speech(text: str, voice: str):
    # Step 1: Gateway decides execution path based on hardware

    decision = gpu_gateway.decide("tts")
    
    # Step 2: Prepare local execution wrapper

    local = gpu_gateway.LocalCall(
        fn=lambda: tts_backend.generate(text, voice),
        what="TTS generation"
    )
    
    # Step 3: Prepare remote fallback (GPU worker cluster)

    remote = gpu_gateway.RemoteCall(
        fn="tts/generate",
        payload={"text": text, "voice": voice, "priority": "high"}
    )
    
    # Step 4: Execute via gateway

    result = await gpu_gateway.run(
        local=local,
        remote=remote,
        decision=decision
    )
    
    # Step 5: Decode audio artifact

    waveform, sample_rate = gpu_gateway.decode_audio_artifact(result)
    return waveform, sample_rate

# Runtime behavior:

# - CUDA/ROCm host: Runs locally on GPU

# - MPS-only Mac: Routes to OmniVoiceMPSSubprocessBackend

# - MLX-capable Mac: Uses MLXAudioBackend

# - No GPU: REMOTE dispatch to GPU worker or CPU fallback

Summary

  • Capability Detection: The system probes hardware using backend/services/model_manager.py with a priority order of CUDA/ROCm > Intel XPU > DirectML > MPS > CPU.
  • Routing Decision: The decide() method in services/gpu_gateway.py selects between LOCAL_CHOSEN, LOCAL_FALLBACK, and REMOTE execution strategies.
  • Backend Abstraction: Concrete classes like OmniVoiceBackend, OmniVoiceMPSSubprocessBackend, and MLXAudioBackend handle platform-specific implementations.
  • ROCm Validation: The gateway verifies gfx architectures to distinguish true ROCm support from CUDA-compatible builds.
  • MPS Constraints: Apple Silicon routing uses a single-worker pool with VRAM budgets derived from system RAM to prevent deadlocks.
  • MLX Availability: MLX backends are exclusively selected on Apple Silicon with Metal support, as enforced by core.device_caps.mlx_supported().

Frequently Asked Questions

How does VoiceStudio handle ROCm devices that report as CUDA-compatible?

VoiceStudio detects ROCm hardware by checking GPU architecture strings (e.g., gfx...) in backend/services/model_manager.py. While ROCm builds report torch.cuda.is_available() == True, the gateway validates the specific architecture against supported ROCm configurations. If validation fails, the system treats the device as CUDA-compatible but without ROCm acceleration, falling back to CPU execution or standard CUDA backends.

What happens when a CUDA-only engine runs on an Apple Silicon Mac?

When gpu_gateway.decide() detects a CUDA-only engine on an MPS-only host, it returns LOCAL_FALLBACK or REMOTE depending on configuration. The LOCAL_FALLBACK state triggers CPU execution, while REMOTE dispatches the request to a GPU worker service running on CUDA-capable hardware via the control-plane scheduler.

Why does MPS routing use a single-worker pool on Apple Silicon?

The MPS backend creates a single-worker pool with HostCaps.vram_gb set to half the system RAM to prevent self-deadlocking patterns. Because MPS shares unified memory between CPU and GPU, concurrent workers could exhaust memory or create circular dependencies. The OmniVoiceMPSSubprocessBackend enforces this constraint to ensure stable inference on macOS devices.

Can MLX backends fall back to remote execution?

No. When core.device_caps.mlx_supported() returns true on Apple Silicon, the gateway selects MLXAudioBackend or MLXWhisperBackend for local execution only. MLX does not support remote dispatch because the framework is exclusive to Apple Silicon architecture. If MLX initialization fails, the gateway falls back to MPS or CPU rather than remote GPU workers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →