VoiceStudio GPU Compatibility Matrix for Long-Form Audiobook and EPUB/PDF Features

VoiceStudio uses a three-layer GPU compatibility matrix—device probe, routing resolver, and UI rendering—to determine whether long-form features like audiobooks and EPUB/PDF import run accelerated, fallback to CPU, or are blocked entirely.

The VoiceStudio open-source TTS platform surfaces GPU compatibility through a deterministic matrix that eliminates silent CPU fallbacks. For resource-intensive long-form pipelines, this matrix guarantees users know upfront whether their audiobook generation or document import will leverage GPU acceleration or face performance degradation. According to the VoiceStudio source code, the matrix is built from device capability detection, engine-specific compatibility declarations, and runtime routing resolution.

How VoiceStudio Detects GPU Capabilities

The foundation of the GPU compatibility matrix is host capability detection, implemented in backend/core/device_caps.py.

The Device Probe Implementation

The detect_host_caps() function lazily imports PyTorch and queries the host for accelerator families, VRAM, driver versions, and edge-case conditions.


# backend/core/device_caps.py (lines 57-92)

from dataclasses import dataclass
from typing import Tuple, Optional, List

@dataclass
class HostCaps:
    family: str                      # 'cuda', 'rocm', 'mps', 'xpu', or 'cpu'

    available_families: Tuple[str, ...]
    vram_gb: Optional[float]
    device_name: Optional[str]
    driver: Optional[str]
    notes: List[str]                 # e.g., ['driver-too-old', 'DirectML']

def detect_host_caps() -> HostCaps:
    """Canonical probe called once per process startup."""
    # Lazily imports torch to avoid heavy dependencies on import

    import torch
    
    # Detection logic for CUDA, ROCm, MPS (Apple Silicon), Intel XPU

    # Returns HostCaps with best available accelerator family

    pass  # See source for full implementation

The probe records critical edge cases: outdated drivers, architecture mismatches, and DirectML presence. These notes flow into the routing decision and UI warnings.

Detected Capabilities

Field Description Example Value
family Best accelerator family "cuda"
available_families All usable families (always includes CPU) ("cuda", "cpu")
vram_gb Available video memory 24.0
device_name GPU product name "NVIDIA GeForce RTX 4090"
driver Driver version string "535.54.03"
notes Diagnostic flags ["driver-too-old"]

Engine GPU Compatibility Declarations

Each TTS and ASR backend declares its supported accelerator families via the gpu_compat tuple. These declarations live in engine definition files and service layer configurations.

TTS Backend Declarations

In backend/services/tts_backend.py at line 172, the VoiceStudio backend declares:


# backend/services/tts_backend.py (line 172)

gpu_compat = ("cuda", "mps", "cpu")

The IndexTTS2 engine, fixed in engines/indextts/__init__.py at line 41, explicitly limits support:


# engines/indextts/__init__.py (line 41)

gpu_compat = ("cuda", "cpu")  # No MPS support

ASR Backend Extensions

ASR backends follow the same pattern in backend/services/asr_backend.py, adding gpu_compat declarations for speech recognition pipelines used in long-form workflows.

Routing Resolution: From Declaration to Effective Device

The routing resolver in backend/services/engine_routing.py intersects engine declarations with host capabilities to produce deterministic outcomes.

The resolve_routing Function


# backend/services/engine_routing.py (lines 33-48)

def resolve_routing(gpu_compat: Tuple[str, ...], caps: HostCaps) -> dict:
    """
    Pure function: no side effects, fully testable.
    Returns: effective_device, routing_status, routing_reason
    """
    if caps.family in gpu_compat and caps.family != "cpu":
        return {
            "effective_device": caps.family,
            "routing_status": "accelerated",
            "routing_reason": None
        }
    elif "cpu" in gpu_compat:
        return {
            "effective_device": "cpu",
            "routing_status": "cpu_fallback" if caps.family != "cpu" else "cpu_only",
            "routing_reason": f"engine has no {caps.family} path; running on CPU"
        }
    else:
        return {
            "effective_device": None,
            "routing_status": "unavailable",
            "routing_reason": f"engine incompatible with {caps.family} and has no CPU fallback"
        }

Routing Status Outcomes

Status Condition User Impact
accelerated Host GPU ∈ gpu_compat Full GPU speed for long-form rendering
cpu_fallback Host GPU ∉ gpu_compat, but cpu is listed 10× slowdown; UI warning displayed
cpu_only Host has no GPU, engine supports CPU Expected CPU operation
unavailable No matching accelerator, no CPU fallback Job blocked with HTTP 400 or WS error
n/a LLM backends (remote-only) Ignored for routing decisions

API and UI Exposure of the Compatibility Matrix

VoiceStudio surfaces the GPU compatibility matrix through multiple interfaces, ensuring visibility before long-form jobs commence.

REST Endpoint: GET /engines

Implemented in backend/api/routers/engines.py, the endpoint returns enriched engine entries:

curl -s https://api.voicestudio.dev/engines | jq '.tts.backends[] | {
  id,
  gpu_compat,
  effective_device,
  routing_status,
  routing_reason
}'

Sample response:

{
  "id": "voice_studio",
  "gpu_compat": ["cuda", "mps", "cpu"],
  "effective_device": "cuda",
  "routing_status": "accelerated",
  "routing_reason": null
}
{
  "id": "vox_cpm2",
  "gpu_compat": ["cuda", "cpu"],
  "effective_device": "cpu",
  "routing_status": "cpu_fallback",
  "routing_reason": "engine has no MPS path; running on CPU"
}

Pre-Flight Wizard Check

The setup wizard (backend/wizard.py) adds a gpu_routing check row that reports the active TTS engine's routing status before any long-form job starts.

Streaming API Headers

Both REST and WebSocket endpoints attach routing metadata:

  • OpenAI-compatible REST (/v1/audio/speech): Headers X-VoiceStudio-Routing and X-VoiceStudio-Routing-Reason
  • WebSocket (/ws/tts): Initial {type: "routing", status, reason} frame

# Python client handling routing headers

import httpx

def synth_longform(payload):
    with httpx.Client(base_url="https://api.voicestudio.dev") as client:
        r = client.post("/longform/render", json=payload, timeout=300)
        routing = r.headers.get("X-VoiceStudio-Routing")
        if routing == "cpu_fallback":
            print("⚠️ Warning:", r.headers.get("X-VoiceStudio-Routing-Reason"))

Frontend GPU Compatibility Matrix UI

The EngineCompatibilityMatrix.jsx component renders the matrix in Settings → Engines, highlighting the effective device and displaying status badges.

Key Rendering Logic

// frontend/src/components/EngineCompatibilityMatrix.jsx (lines 28-55)
const CHIP_EFFECTIVE = 'is-effective';

const ROUTING_BADGE = {
  accelerated:  { tone: 'success', labelKey: 'engines.routingAccelerated' },
  cpu_fallback: { tone: 'warn',    labelKey: 'engines.routingCpuFallback' },
  cpu_only:     { tone: 'neutral', labelKey: 'engines.routingCpuOnly' },
  'n/a':        { tone: 'neutral', labelKey: 'engines.routingRemote' },
};

function normalizeEntry(engineData) {
  const {
    effective_device,
    routing_status,
    routing_reason
  } = engineData;
  
  return {
    effectiveDevice: effective_device,
    status: routing_status,
    reason: routing_reason,
    badge: ROUTING_BADGE[routing_status]
  };
}

The component applies CHIP_EFFECTIVE to the GPU chip matching effective_device, ensuring users instantly see which accelerator is active.

Client-Side Routing Checks Before Long-Form Jobs

Implement proactive compatibility checks to warn users before starting audiobook or EPUB/PDF conversion:

// Client-side check before launching long-form generation
async function canRunLongform(engineId) {
  const resp = await fetch('/engines');
  const data = await resp.json();
  const engine = data.tts.backends.find(e => e.id === engineId);
  if (!engine) throw new Error('Engine not found');

  switch (engine.routing_status) {
    case 'accelerated':
      return true; // Optimal path
    case 'cpu_fallback':
      console.warn(engine.routing_reason);
      return confirm('GPU unavailable. Run on CPU (significantly slower)?');
    case 'unavailable':
      alert(`Cannot run on this host: ${engine.routing_reason}`);
      return false;
    default:
      return false;
  }
}

Why GPU Compatibility Matters for Long-Form Features

Long-form rendering (/longform/render, /audiobook) streams audio chunks through the GPU-pool worker. Silent CPU fallbacks cause 10× performance degradation on hour-long audiobooks or multi-chapter EPUB conversions.

The VoiceStudio GPU compatibility matrix prevents this by:

  • Explicit warnings: cpu_fallback triggers UI badges and pre-flight warnings
  • Hard blocking: unavailable returns HTTP 400 or WebSocket errors rather than failing mid-job
  • Performance transparency: Device name and VRAM display set latency expectations

Summary

  • Device probe (backend/core/device_caps.py): Detects CUDA, ROCm, MPS, XPU, or CPU with driver/VRAM diagnostics
  • Engine declarations (backend/services/tts_backend.py, engine modules): Each backend lists compatible GPU families in gpu_compat
  • Routing resolver (backend/services/engine_routing.py): Pure function produces effective_device, routing_status, and routing_reason
  • Five routing statuses: accelerated, cpu_fallback, cpu_only, unavailable, n/a for LLMs
  • Multi-interface exposure: REST /engines, wizard pre-flight, HTTP headers, WebSocket frames, and React UI matrix
  • Long-form protection: Eliminates silent CPU fallbacks that would degrade audiobook and EPUB/PDF import performance by 10×

Frequently Asked Questions

What GPU families does VoiceStudio support for long-form features?

VoiceStudio supports CUDA (NVIDIA), ROCm (AMD), MPS (Apple Silicon), XPU (Intel), and CPU fallback. The device probe in backend/core/device_caps.py detects all five families, with family indicating the best available accelerator and available_families listing all usable options.

How do I check if my audiobook job will run on GPU before starting?

Query the /engines endpoint and inspect routing_status for your selected TTS backend. Status accelerated confirms GPU execution. Status cpu_fallback warns of CPU degradation. The pre-flight wizard (backend/wizard.py) and streaming API headers provide the same information without manual API calls.

Why does VoiceStudio block some engines instead of falling back to CPU?

Engines with routing_status: unavailable lack both the host's GPU family and cpu in their gpu_compat declaration. VoiceStudio blocks these rather than fail mid-job. This occurs when specialized engines require specific hardware (e.g., CUDA-only models on Apple Silicon hosts without CPU support declared).

Can I extend GPU compatibility to custom TTS backends?

Yes. Add a gpu_compat tuple to your engine's initialization file (pattern: engines/your_engine/__init__.py). Include "cuda", "rocm", "mps", "xpu", and/or "cpu" based on your model's tested platforms. The routing resolver automatically incorporates your engine into the compatibility matrix without frontend changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →