# VoiceStudio GPU Compatibility Matrix for Long-Form Audiobook and EPUB/PDF Features

> Explore the VoiceStudio GPU compatibility matrix for accelerated audiobook and EPUB PDF import. Understand device probe, routing, and UI rendering for your features.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: getting-started
- Published: 2026-09-06

---

**VoiceStudio uses a three-layer GPU compatibility matrix—device probe, routing resolver, and UI rendering—to determine whether long-form features like audiobooks and EPUB/PDF import run accelerated, fallback to CPU, or are blocked entirely.**

The VoiceStudio open-source TTS platform surfaces GPU compatibility through a deterministic matrix that eliminates silent CPU fallbacks. For resource-intensive long-form pipelines, this matrix guarantees users know upfront whether their audiobook generation or document import will leverage GPU acceleration or face performance degradation. According to the VoiceStudio source code, the matrix is built from device capability detection, engine-specific compatibility declarations, and runtime routing resolution.

## How VoiceStudio Detects GPU Capabilities

The foundation of the GPU compatibility matrix is **host capability detection**, implemented in [`backend/core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/device_caps.py).

### The Device Probe Implementation

The `detect_host_caps()` function lazily imports PyTorch and queries the host for accelerator families, VRAM, driver versions, and edge-case conditions.

```python

# backend/core/device_caps.py (lines 57-92)

from dataclasses import dataclass
from typing import Tuple, Optional, List

@dataclass
class HostCaps:
    family: str                      # 'cuda', 'rocm', 'mps', 'xpu', or 'cpu'

    available_families: Tuple[str, ...]
    vram_gb: Optional[float]
    device_name: Optional[str]
    driver: Optional[str]
    notes: List[str]                 # e.g., ['driver-too-old', 'DirectML']

def detect_host_caps() -> HostCaps:
    """Canonical probe called once per process startup."""
    # Lazily imports torch to avoid heavy dependencies on import

    import torch
    
    # Detection logic for CUDA, ROCm, MPS (Apple Silicon), Intel XPU

    # Returns HostCaps with best available accelerator family

    pass  # See source for full implementation

```

The probe records critical edge cases: outdated drivers, architecture mismatches, and DirectML presence. These **notes** flow into the routing decision and UI warnings.

### Detected Capabilities

| Field | Description | Example Value |
|-------|-------------|---------------|
| `family` | Best accelerator family | `"cuda"` |
| `available_families` | All usable families (always includes CPU) | `("cuda", "cpu")` |
| `vram_gb` | Available video memory | `24.0` |
| `device_name` | GPU product name | `"NVIDIA GeForce RTX 4090"` |
| `driver` | Driver version string | `"535.54.03"` |
| `notes` | Diagnostic flags | `["driver-too-old"]` |

## Engine GPU Compatibility Declarations

Each TTS and ASR backend declares its supported accelerator families via the `gpu_compat` tuple. These declarations live in engine definition files and service layer configurations.

### TTS Backend Declarations

In [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py) at line 172, the VoiceStudio backend declares:

```python

# backend/services/tts_backend.py (line 172)

gpu_compat = ("cuda", "mps", "cpu")

```

The IndexTTS2 engine, fixed in [`engines/indextts/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/engines/indextts/__init__.py) at line 41, explicitly limits support:

```python

# engines/indextts/__init__.py (line 41)

gpu_compat = ("cuda", "cpu")  # No MPS support

```

### ASR Backend Extensions

ASR backends follow the same pattern in [`backend/services/asr_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/asr_backend.py), adding `gpu_compat` declarations for speech recognition pipelines used in long-form workflows.

## Routing Resolution: From Declaration to Effective Device

The **routing resolver** in [`backend/services/engine_routing.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py) intersects engine declarations with host capabilities to produce deterministic outcomes.

### The resolve_routing Function

```python

# backend/services/engine_routing.py (lines 33-48)

def resolve_routing(gpu_compat: Tuple[str, ...], caps: HostCaps) -> dict:
    """
    Pure function: no side effects, fully testable.
    Returns: effective_device, routing_status, routing_reason
    """
    if caps.family in gpu_compat and caps.family != "cpu":
        return {
            "effective_device": caps.family,
            "routing_status": "accelerated",
            "routing_reason": None
        }
    elif "cpu" in gpu_compat:
        return {
            "effective_device": "cpu",
            "routing_status": "cpu_fallback" if caps.family != "cpu" else "cpu_only",
            "routing_reason": f"engine has no {caps.family} path; running on CPU"
        }
    else:
        return {
            "effective_device": None,
            "routing_status": "unavailable",
            "routing_reason": f"engine incompatible with {caps.family} and has no CPU fallback"
        }

```

### Routing Status Outcomes

| Status | Condition | User Impact |
|--------|-----------|-------------|
| **accelerated** | Host GPU ∈ `gpu_compat` | Full GPU speed for long-form rendering |
| **cpu_fallback** | Host GPU ∉ `gpu_compat`, but `cpu` is listed | 10× slowdown; UI warning displayed |
| **cpu_only** | Host has no GPU, engine supports CPU | Expected CPU operation |
| **unavailable** | No matching accelerator, no CPU fallback | Job blocked with HTTP 400 or WS error |
| **n/a** | LLM backends (remote-only) | Ignored for routing decisions |

## API and UI Exposure of the Compatibility Matrix

VoiceStudio surfaces the GPU compatibility matrix through multiple interfaces, ensuring visibility before long-form jobs commence.

### REST Endpoint: GET /engines

Implemented in [`backend/api/routers/engines.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/engines.py), the endpoint returns enriched engine entries:

```bash
curl -s https://api.voicestudio.dev/engines | jq '.tts.backends[] | {
  id,
  gpu_compat,
  effective_device,
  routing_status,
  routing_reason
}'

```

**Sample response:**

```json
{
  "id": "voice_studio",
  "gpu_compat": ["cuda", "mps", "cpu"],
  "effective_device": "cuda",
  "routing_status": "accelerated",
  "routing_reason": null
}
{
  "id": "vox_cpm2",
  "gpu_compat": ["cuda", "cpu"],
  "effective_device": "cpu",
  "routing_status": "cpu_fallback",
  "routing_reason": "engine has no MPS path; running on CPU"
}

```

### Pre-Flight Wizard Check

The setup wizard ([`backend/wizard.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/wizard.py)) adds a `gpu_routing` check row that reports the active TTS engine's routing status before any long-form job starts.

### Streaming API Headers

Both REST and WebSocket endpoints attach routing metadata:

- **OpenAI-compatible REST** (`/v1/audio/speech`): Headers `X-VoiceStudio-Routing` and `X-VoiceStudio-Routing-Reason`
- **WebSocket** (`/ws/tts`): Initial `{type: "routing", status, reason}` frame

```python

# Python client handling routing headers

import httpx

def synth_longform(payload):
    with httpx.Client(base_url="https://api.voicestudio.dev") as client:
        r = client.post("/longform/render", json=payload, timeout=300)
        routing = r.headers.get("X-VoiceStudio-Routing")
        if routing == "cpu_fallback":
            print("⚠️ Warning:", r.headers.get("X-VoiceStudio-Routing-Reason"))

```

## Frontend GPU Compatibility Matrix UI

The [`EngineCompatibilityMatrix.jsx`](https://github.com/debpalash/VoiceStudio/blob/main/EngineCompatibilityMatrix.jsx) component renders the matrix in Settings → Engines, highlighting the effective device and displaying status badges.

### Key Rendering Logic

```jsx
// frontend/src/components/EngineCompatibilityMatrix.jsx (lines 28-55)
const CHIP_EFFECTIVE = 'is-effective';

const ROUTING_BADGE = {
  accelerated:  { tone: 'success', labelKey: 'engines.routingAccelerated' },
  cpu_fallback: { tone: 'warn',    labelKey: 'engines.routingCpuFallback' },
  cpu_only:     { tone: 'neutral', labelKey: 'engines.routingCpuOnly' },
  'n/a':        { tone: 'neutral', labelKey: 'engines.routingRemote' },
};

function normalizeEntry(engineData) {
  const {
    effective_device,
    routing_status,
    routing_reason
  } = engineData;
  
  return {
    effectiveDevice: effective_device,
    status: routing_status,
    reason: routing_reason,
    badge: ROUTING_BADGE[routing_status]
  };
}

```

The component applies `CHIP_EFFECTIVE` to the GPU chip matching `effective_device`, ensuring users instantly see which accelerator is active.

## Client-Side Routing Checks Before Long-Form Jobs

Implement proactive compatibility checks to warn users before starting audiobook or EPUB/PDF conversion:

```javascript
// Client-side check before launching long-form generation
async function canRunLongform(engineId) {
  const resp = await fetch('/engines');
  const data = await resp.json();
  const engine = data.tts.backends.find(e => e.id === engineId);
  if (!engine) throw new Error('Engine not found');

  switch (engine.routing_status) {
    case 'accelerated':
      return true; // Optimal path
    case 'cpu_fallback':
      console.warn(engine.routing_reason);
      return confirm('GPU unavailable. Run on CPU (significantly slower)?');
    case 'unavailable':
      alert(`Cannot run on this host: ${engine.routing_reason}`);
      return false;
    default:
      return false;
  }
}

```

## Why GPU Compatibility Matters for Long-Form Features

Long-form rendering (`/longform/render`, `/audiobook`) streams audio chunks through the **GPU-pool worker**. Silent CPU fallbacks cause **10× performance degradation** on hour-long audiobooks or multi-chapter EPUB conversions.

The VoiceStudio GPU compatibility matrix prevents this by:

- **Explicit warnings**: `cpu_fallback` triggers UI badges and pre-flight warnings
- **Hard blocking**: `unavailable` returns HTTP 400 or WebSocket errors rather than failing mid-job
- **Performance transparency**: Device name and VRAM display set latency expectations

## Summary

- **Device probe** ([`backend/core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/device_caps.py)): Detects CUDA, ROCm, MPS, XPU, or CPU with driver/VRAM diagnostics
- **Engine declarations** ([`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py), engine modules): Each backend lists compatible GPU families in `gpu_compat`
- **Routing resolver** ([`backend/services/engine_routing.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/engine_routing.py)): Pure function produces `effective_device`, `routing_status`, and `routing_reason`
- **Five routing statuses**: `accelerated`, `cpu_fallback`, `cpu_only`, `unavailable`, `n/a` for LLMs
- **Multi-interface exposure**: REST `/engines`, wizard pre-flight, HTTP headers, WebSocket frames, and React UI matrix
- **Long-form protection**: Eliminates silent CPU fallbacks that would degrade audiobook and EPUB/PDF import performance by 10×

## Frequently Asked Questions

### What GPU families does VoiceStudio support for long-form features?

VoiceStudio supports **CUDA** (NVIDIA), **ROCm** (AMD), **MPS** (Apple Silicon), **XPU** (Intel), and **CPU** fallback. The device probe in [`backend/core/device_caps.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/core/device_caps.py) detects all five families, with `family` indicating the best available accelerator and `available_families` listing all usable options.

### How do I check if my audiobook job will run on GPU before starting?

Query the `/engines` endpoint and inspect `routing_status` for your selected TTS backend. Status `accelerated` confirms GPU execution. Status `cpu_fallback` warns of CPU degradation. The pre-flight wizard ([`backend/wizard.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/wizard.py)) and streaming API headers provide the same information without manual API calls.

### Why does VoiceStudio block some engines instead of falling back to CPU?

Engines with `routing_status: unavailable` lack both the host's GPU family and `cpu` in their `gpu_compat` declaration. VoiceStudio blocks these rather than fail mid-job. This occurs when specialized engines require specific hardware (e.g., CUDA-only models on Apple Silicon hosts without CPU support declared).

### Can I extend GPU compatibility to custom TTS backends?

Yes. Add a `gpu_compat` tuple to your engine's initialization file (pattern: [`engines/your_engine/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/engines/your_engine/__init__.py)). Include `"cuda"`, `"rocm"`, `"mps"`, `"xpu"`, and/or `"cpu"` based on your model's tested platforms. The routing resolver automatically incorporates your engine into the compatibility matrix without frontend changes.