# How VoiceStudio's GPU Gateway Handles CUDA, MPS, and ROCm Engine Constraints

> Discover how VoiceStudio's GPU gateway manages CUDA, MPS, and ROCm constraints, ensuring efficient job execution by matching workloads with compatible pools and adjusting timeouts.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: architecture
- Published: 2026-09-12

---

**VoiceStudio's GPU gateway enforces hardware-specific constraints by admitting jobs only when compatible CUDA, MPS, or ROCm pools have capacity, adjusting execution timeouts per device family, and delegating to remote workers when local accelerators cannot satisfy engine requirements.**

VoiceStudio routes every GPU-accelerated inference request through [`backend/services/gpu_gateway.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/gpu_gateway.py), a centralized entry point that guarantees safe execution on compatible hardware. Rather than guessing device capabilities at runtime, the gateway enforces pre-computed routing decisions that account for CUDA, MPS, and ROCm availability while managing pool admission and graceful fallbacks.

## Device Detection and Dynamic Worker Pool Sizing

The groundwork for engine constraint handling begins in [`backend/services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/model_manager.py), where the system probes the host's hardware capabilities before the gateway ever receives a request.

### Detecting CUDA, MPS, and ROCm Availability

The `model_manager._pick_gpu_workers` method probes the PyTorch runtime to determine which acceleration family is present. It checks `torch.cuda.is_available()` for NVIDIA hardware, `torch.backends.mps.is_available()` for Apple Silicon, and equivalent ROCm APIs, then constructs an appropriately sized worker pool based on the detected family.

### Pool Sizing Rules per Engine Family

The sizing logic respects the architectural limits of each platform:

- **CUDA and ROCm**: Pool size equals free VRAM divided by the per-job budget, capped at four workers to prevent oversubscription.
- **MPS**: Hard-limited to a single worker because the memory pool is shared with system RAM, as enforced at `model_manager.py#L95-L98`.

This sizing occurs before the gateway accepts traffic, ensuring that the admission controller knows the exact capacity of each hardware family.

## Routing Decisions and Constraint Enforcement

The gateway itself does not decide which hardware family can service a request; instead, it relies on the device-capability subsystem and routing layer to pre-calculate feasibility.

### Pre-Computing the Execution Target

When a request arrives, `gpu_gateway.decide` delegates to `worker.routing.decide`, which consults [`core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/device_caps.py) and [`services/engine_routing.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/engine_routing.py) to match the engine's requirements against advertised capabilities. The result is a `Decision` object that tells the gateway whether to execute locally on a CUDA/MPS/ROCm device or forward the job to a remote worker.

According to the source at `gpu_gateway.py#L4-L13`, this decision encapsulates the engine constraints discovered earlier, preventing the gateway from attempting to run a CUDA-specific model on a CPU-only host.

### Local vs. Remote GPU Handling

The gateway respects the decision's directive by switching between two distinct code paths:

- **Local execution**: Uses `gpu_gateway._run_local`, which performs admission checks against the local GPU pool.
- **Remote execution**: Skips local pool interactions entirely, forwarding the request via `gpu_gateway._run_remote` to a worker that advertises the required CUDA, ROCm, or MPS capabilities.

This separation prevents the gateway from exhausting local resources on engines the hardware cannot support.

## Admission Control and Timeout Management

Once the gateway determines local execution is viable, it enforces runtime constraints through admission control and resource-aware timeouts.

### GPU Pool Admission Checks

Before submitting work to a local device, `gpu_gateway._run_local` executes `check_gpu_admission`. This function reads `gpu_pool_stats` and raises `GpuPoolBusyError` if the queue is saturated, as implemented at `gpu_gateway.py#L46-L53`. This check ensures that a CUDA or ROCm device already running at capacity refuses new work rather than accepting jobs it cannot complete.

### Engine-Aware Timeout Selection

The `model_manager.generate_timeout_s` method tailors execution budgets to the detected hardware family. If the effective family is `"cpu"`, the system applies the CPU timeout; otherwise, it uses the accelerated timeout. However, under-provisioned GPUs—including low-VRAM CUDA/ROCm cards or MPS devices—automatically fall back to the CPU timeout to avoid premature cancellation, as shown at `model_manager.py#L89-L96`.

## MPS-Specific Safeguards

Apple Silicon's MPS backend requires special handling because it shares system memory rather than using dedicated VRAM. The gateway enforces two critical safeguards:

1. **Single-worker limitation**: The `_pick_gpu_workers` function logs "MPS detected, using 1 worker (shared system memory)" and caps the pool at one worker (`model_manager.py#L95-L98`).
2. **Deadlock prevention**: All code paths referencing `running_on_gpu_pool()` treat MPS as a single-worker environment, preventing the system from submitting concurrent jobs that would self-deadlock on the one-worker pool (`model_manager.py#L16-L23`).

## Remote-First Execution and Fallback Patterns

When the routing decision indicates remote execution, the gateway skips local pool prewarming and immediately delegates to `gpu_gateway._run_remote`. This avoids attempting to initialize CUDA or MPS contexts on machines that lack the hardware, while the remote worker independently validates its own capabilities before accepting the job.

For resilience, the `JobRun` object tracks consecutive remote failures and automatically latches to local execution after crossing the configured `_MULTI_UNIT_FAILURE_LIMIT`, ensuring that a transient network partition does not permanently disable GPU acceleration.

## Implementation Example

The following pattern demonstrates how client code interacts with the gateway to respect engine constraints:

```python

# 1. Query the routing layer for a feasible target (local GPU, remote GPU, or CPU)

decision = gpu_gateway.decide("tts")  # Consults device_caps and routing logic

# 2. Prewarm the chosen backend (noop if remote)

await gpu_gateway.prewarm(
    "tts",
    backend=tts_backend,
    engine="omnivore",
    decision=decision,
)

# 3. Execute, letting the gateway handle admission and fallback

audio = await gpu_gateway.run(
    "tts",
    local=gpu_gateway.LocalCall(fn=tts_backend.generate),
    remote=gpu_gateway.RemoteCall(
        engine="omnivore",
        params={"text": "Hello world"},
        operation="tts",
        decode=gpu_gateway.decode_audio_artifact,
    ),
    decision=decision,
)

```

This flow works identically for batch jobs; the gateway transparently handles whether the underlying engine requires CUDA, ROCm, or MPS without the caller managing hardware detection.

## Summary

- **Pre-computed decisions**: The gateway relies on [`worker/routing.py`](https://github.com/debpalash/VoiceStudio/blob/main/worker/routing.py) and [`core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/core/device_caps.py) to determine CUDA, MPS, and ROCm feasibility before admission.
- **Pool-aware admission**: `gpu_gateway._run_local` rejects jobs via `check_gpu_admission` when the local GPU pool is saturated (`gpu_gateway.py#L46-L53`).
- **Dynamic timeouts**: `model_manager.generate_timeout_s` assigns CPU-grade timeouts to MPS and low-VRAM devices to prevent false-positive timeouts (`model_manager.py#L89-L96`).
- **MPS isolation**: Single-worker pools prevent memory contention and deadlock on Apple Silicon (`model_manager.py#L95-L98`).
- **Remote delegation**: When local constraints cannot be satisfied, the gateway transparently forwards work to compatible remote workers.

## Frequently Asked Questions

### What happens if the local machine lacks the GPU family required by an engine?

The routing layer returns a remote-only `Decision`, causing `gpu_gateway.run` to skip local pool admission entirely and invoke `gpu_gateway._run_remote`. The remote worker then validates its own CUDA, ROCm, or MPS capabilities before accepting the job, ensuring the engine never executes on incompatible hardware.

### How does VoiceStudio prevent MPS memory exhaustion on Apple Silicon?

The system caps the MPS worker pool at a single thread because MPS shares system RAM rather than using dedicated VRAM. This is enforced in `model_manager._pick_gpu_workers` (`model_manager.py#L95-L98`) and respected by all `running_on_gpu_pool()` checks to prevent concurrent access that could deadlock or exhaust memory.

### Why do low-VRAM CUDA cards receive CPU timeouts instead of GPU timeouts?

The `generate_timeout_s` function identifies under-provisioned GPUs—including low-VRAM CUDA/ROCm devices and MPS—and falls back to the CPU timeout budget (`model_manager.py#L89-L96`). This prevents the scheduler from assuming accelerated performance when the hardware cannot deliver it, avoiding premature job cancellation.

### Can the gateway mix different engine families in the same request batch?

Yes, but each sub-job is routed independently. The `JobRun` object tracks consecutive failures per unit and switches to local execution after exceeding `_MULTI_UNIT_FAILURE_LIMIT`, but the routing decision for each engine is evaluated separately against the local device's CUDA, MPS, or ROCm capabilities.