How VoiceStudio's GPU Gateway Handles CUDA, MPS, and ROCm Engine Constraints

VoiceStudio's GPU gateway enforces hardware-specific constraints by admitting jobs only when compatible CUDA, MPS, or ROCm pools have capacity, adjusting execution timeouts per device family, and delegating to remote workers when local accelerators cannot satisfy engine requirements.

VoiceStudio routes every GPU-accelerated inference request through backend/services/gpu_gateway.py, a centralized entry point that guarantees safe execution on compatible hardware. Rather than guessing device capabilities at runtime, the gateway enforces pre-computed routing decisions that account for CUDA, MPS, and ROCm availability while managing pool admission and graceful fallbacks.

Device Detection and Dynamic Worker Pool Sizing

The groundwork for engine constraint handling begins in backend/services/model_manager.py, where the system probes the host's hardware capabilities before the gateway ever receives a request.

Detecting CUDA, MPS, and ROCm Availability

The model_manager._pick_gpu_workers method probes the PyTorch runtime to determine which acceleration family is present. It checks torch.cuda.is_available() for NVIDIA hardware, torch.backends.mps.is_available() for Apple Silicon, and equivalent ROCm APIs, then constructs an appropriately sized worker pool based on the detected family.

Pool Sizing Rules per Engine Family

The sizing logic respects the architectural limits of each platform:

  • CUDA and ROCm: Pool size equals free VRAM divided by the per-job budget, capped at four workers to prevent oversubscription.
  • MPS: Hard-limited to a single worker because the memory pool is shared with system RAM, as enforced at model_manager.py#L95-L98.

This sizing occurs before the gateway accepts traffic, ensuring that the admission controller knows the exact capacity of each hardware family.

Routing Decisions and Constraint Enforcement

The gateway itself does not decide which hardware family can service a request; instead, it relies on the device-capability subsystem and routing layer to pre-calculate feasibility.

Pre-Computing the Execution Target

When a request arrives, gpu_gateway.decide delegates to worker.routing.decide, which consults core/device_caps.py and services/engine_routing.py to match the engine's requirements against advertised capabilities. The result is a Decision object that tells the gateway whether to execute locally on a CUDA/MPS/ROCm device or forward the job to a remote worker.

According to the source at gpu_gateway.py#L4-L13, this decision encapsulates the engine constraints discovered earlier, preventing the gateway from attempting to run a CUDA-specific model on a CPU-only host.

Local vs. Remote GPU Handling

The gateway respects the decision's directive by switching between two distinct code paths:

  • Local execution: Uses gpu_gateway._run_local, which performs admission checks against the local GPU pool.
  • Remote execution: Skips local pool interactions entirely, forwarding the request via gpu_gateway._run_remote to a worker that advertises the required CUDA, ROCm, or MPS capabilities.

This separation prevents the gateway from exhausting local resources on engines the hardware cannot support.

Admission Control and Timeout Management

Once the gateway determines local execution is viable, it enforces runtime constraints through admission control and resource-aware timeouts.

GPU Pool Admission Checks

Before submitting work to a local device, gpu_gateway._run_local executes check_gpu_admission. This function reads gpu_pool_stats and raises GpuPoolBusyError if the queue is saturated, as implemented at gpu_gateway.py#L46-L53. This check ensures that a CUDA or ROCm device already running at capacity refuses new work rather than accepting jobs it cannot complete.

Engine-Aware Timeout Selection

The model_manager.generate_timeout_s method tailors execution budgets to the detected hardware family. If the effective family is "cpu", the system applies the CPU timeout; otherwise, it uses the accelerated timeout. However, under-provisioned GPUs—including low-VRAM CUDA/ROCm cards or MPS devices—automatically fall back to the CPU timeout to avoid premature cancellation, as shown at model_manager.py#L89-L96.

MPS-Specific Safeguards

Apple Silicon's MPS backend requires special handling because it shares system memory rather than using dedicated VRAM. The gateway enforces two critical safeguards:

  1. Single-worker limitation: The _pick_gpu_workers function logs "MPS detected, using 1 worker (shared system memory)" and caps the pool at one worker (model_manager.py#L95-L98).
  2. Deadlock prevention: All code paths referencing running_on_gpu_pool() treat MPS as a single-worker environment, preventing the system from submitting concurrent jobs that would self-deadlock on the one-worker pool (model_manager.py#L16-L23).

Remote-First Execution and Fallback Patterns

When the routing decision indicates remote execution, the gateway skips local pool prewarming and immediately delegates to gpu_gateway._run_remote. This avoids attempting to initialize CUDA or MPS contexts on machines that lack the hardware, while the remote worker independently validates its own capabilities before accepting the job.

For resilience, the JobRun object tracks consecutive remote failures and automatically latches to local execution after crossing the configured _MULTI_UNIT_FAILURE_LIMIT, ensuring that a transient network partition does not permanently disable GPU acceleration.

Implementation Example

The following pattern demonstrates how client code interacts with the gateway to respect engine constraints:


# 1. Query the routing layer for a feasible target (local GPU, remote GPU, or CPU)

decision = gpu_gateway.decide("tts")  # Consults device_caps and routing logic

# 2. Prewarm the chosen backend (noop if remote)

await gpu_gateway.prewarm(
    "tts",
    backend=tts_backend,
    engine="omnivore",
    decision=decision,
)

# 3. Execute, letting the gateway handle admission and fallback

audio = await gpu_gateway.run(
    "tts",
    local=gpu_gateway.LocalCall(fn=tts_backend.generate),
    remote=gpu_gateway.RemoteCall(
        engine="omnivore",
        params={"text": "Hello world"},
        operation="tts",
        decode=gpu_gateway.decode_audio_artifact,
    ),
    decision=decision,
)

This flow works identically for batch jobs; the gateway transparently handles whether the underlying engine requires CUDA, ROCm, or MPS without the caller managing hardware detection.

Summary

  • Pre-computed decisions: The gateway relies on worker/routing.py and core/device_caps.py to determine CUDA, MPS, and ROCm feasibility before admission.
  • Pool-aware admission: gpu_gateway._run_local rejects jobs via check_gpu_admission when the local GPU pool is saturated (gpu_gateway.py#L46-L53).
  • Dynamic timeouts: model_manager.generate_timeout_s assigns CPU-grade timeouts to MPS and low-VRAM devices to prevent false-positive timeouts (model_manager.py#L89-L96).
  • MPS isolation: Single-worker pools prevent memory contention and deadlock on Apple Silicon (model_manager.py#L95-L98).
  • Remote delegation: When local constraints cannot be satisfied, the gateway transparently forwards work to compatible remote workers.

Frequently Asked Questions

What happens if the local machine lacks the GPU family required by an engine?

The routing layer returns a remote-only Decision, causing gpu_gateway.run to skip local pool admission entirely and invoke gpu_gateway._run_remote. The remote worker then validates its own CUDA, ROCm, or MPS capabilities before accepting the job, ensuring the engine never executes on incompatible hardware.

How does VoiceStudio prevent MPS memory exhaustion on Apple Silicon?

The system caps the MPS worker pool at a single thread because MPS shares system RAM rather than using dedicated VRAM. This is enforced in model_manager._pick_gpu_workers (model_manager.py#L95-L98) and respected by all running_on_gpu_pool() checks to prevent concurrent access that could deadlock or exhaust memory.

Why do low-VRAM CUDA cards receive CPU timeouts instead of GPU timeouts?

The generate_timeout_s function identifies under-provisioned GPUs—including low-VRAM CUDA/ROCm devices and MPS—and falls back to the CPU timeout budget (model_manager.py#L89-L96). This prevents the scheduler from assuming accelerated performance when the hardware cannot deliver it, avoiding premature job cancellation.

Can the gateway mix different engine families in the same request batch?

Yes, but each sub-job is routed independently. The JobRun object tracks consecutive failures per unit and switches to local execution after exceeding _MULTI_UNIT_FAILURE_LIMIT, but the routing decision for each engine is evaluated separately against the local device's CUDA, MPS, or ROCm capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →