VoiceStudio GPU Compatibility Matrix for Long-Form Audiobook and EPUB/PDF Features
VoiceStudio uses a three-layer GPU compatibility matrix—device probe, routing resolver, and UI rendering—to determine whether long-form features like audiobooks and EPUB/PDF import run accelerated, fallback to CPU, or are blocked entirely.
The VoiceStudio open-source TTS platform surfaces GPU compatibility through a deterministic matrix that eliminates silent CPU fallbacks. For resource-intensive long-form pipelines, this matrix guarantees users know upfront whether their audiobook generation or document import will leverage GPU acceleration or face performance degradation. According to the VoiceStudio source code, the matrix is built from device capability detection, engine-specific compatibility declarations, and runtime routing resolution.
How VoiceStudio Detects GPU Capabilities
The foundation of the GPU compatibility matrix is host capability detection, implemented in backend/core/device_caps.py.
The Device Probe Implementation
The detect_host_caps() function lazily imports PyTorch and queries the host for accelerator families, VRAM, driver versions, and edge-case conditions.
# backend/core/device_caps.py (lines 57-92)
from dataclasses import dataclass
from typing import Tuple, Optional, List
@dataclass
class HostCaps:
family: str # 'cuda', 'rocm', 'mps', 'xpu', or 'cpu'
available_families: Tuple[str, ...]
vram_gb: Optional[float]
device_name: Optional[str]
driver: Optional[str]
notes: List[str] # e.g., ['driver-too-old', 'DirectML']
def detect_host_caps() -> HostCaps:
"""Canonical probe called once per process startup."""
# Lazily imports torch to avoid heavy dependencies on import
import torch
# Detection logic for CUDA, ROCm, MPS (Apple Silicon), Intel XPU
# Returns HostCaps with best available accelerator family
pass # See source for full implementation
The probe records critical edge cases: outdated drivers, architecture mismatches, and DirectML presence. These notes flow into the routing decision and UI warnings.
Detected Capabilities
| Field | Description | Example Value |
|---|---|---|
family |
Best accelerator family | "cuda" |
available_families |
All usable families (always includes CPU) | ("cuda", "cpu") |
vram_gb |
Available video memory | 24.0 |
device_name |
GPU product name | "NVIDIA GeForce RTX 4090" |
driver |
Driver version string | "535.54.03" |
notes |
Diagnostic flags | ["driver-too-old"] |
Engine GPU Compatibility Declarations
Each TTS and ASR backend declares its supported accelerator families via the gpu_compat tuple. These declarations live in engine definition files and service layer configurations.
TTS Backend Declarations
In backend/services/tts_backend.py at line 172, the VoiceStudio backend declares:
# backend/services/tts_backend.py (line 172)
gpu_compat = ("cuda", "mps", "cpu")
The IndexTTS2 engine, fixed in engines/indextts/__init__.py at line 41, explicitly limits support:
# engines/indextts/__init__.py (line 41)
gpu_compat = ("cuda", "cpu") # No MPS support
ASR Backend Extensions
ASR backends follow the same pattern in backend/services/asr_backend.py, adding gpu_compat declarations for speech recognition pipelines used in long-form workflows.
Routing Resolution: From Declaration to Effective Device
The routing resolver in backend/services/engine_routing.py intersects engine declarations with host capabilities to produce deterministic outcomes.
The resolve_routing Function
# backend/services/engine_routing.py (lines 33-48)
def resolve_routing(gpu_compat: Tuple[str, ...], caps: HostCaps) -> dict:
"""
Pure function: no side effects, fully testable.
Returns: effective_device, routing_status, routing_reason
"""
if caps.family in gpu_compat and caps.family != "cpu":
return {
"effective_device": caps.family,
"routing_status": "accelerated",
"routing_reason": None
}
elif "cpu" in gpu_compat:
return {
"effective_device": "cpu",
"routing_status": "cpu_fallback" if caps.family != "cpu" else "cpu_only",
"routing_reason": f"engine has no {caps.family} path; running on CPU"
}
else:
return {
"effective_device": None,
"routing_status": "unavailable",
"routing_reason": f"engine incompatible with {caps.family} and has no CPU fallback"
}
Routing Status Outcomes
| Status | Condition | User Impact |
|---|---|---|
| accelerated | Host GPU ∈ gpu_compat |
Full GPU speed for long-form rendering |
| cpu_fallback | Host GPU ∉ gpu_compat, but cpu is listed |
10× slowdown; UI warning displayed |
| cpu_only | Host has no GPU, engine supports CPU | Expected CPU operation |
| unavailable | No matching accelerator, no CPU fallback | Job blocked with HTTP 400 or WS error |
| n/a | LLM backends (remote-only) | Ignored for routing decisions |
API and UI Exposure of the Compatibility Matrix
VoiceStudio surfaces the GPU compatibility matrix through multiple interfaces, ensuring visibility before long-form jobs commence.
REST Endpoint: GET /engines
Implemented in backend/api/routers/engines.py, the endpoint returns enriched engine entries:
curl -s https://api.voicestudio.dev/engines | jq '.tts.backends[] | {
id,
gpu_compat,
effective_device,
routing_status,
routing_reason
}'
Sample response:
{
"id": "voice_studio",
"gpu_compat": ["cuda", "mps", "cpu"],
"effective_device": "cuda",
"routing_status": "accelerated",
"routing_reason": null
}
{
"id": "vox_cpm2",
"gpu_compat": ["cuda", "cpu"],
"effective_device": "cpu",
"routing_status": "cpu_fallback",
"routing_reason": "engine has no MPS path; running on CPU"
}
Pre-Flight Wizard Check
The setup wizard (backend/wizard.py) adds a gpu_routing check row that reports the active TTS engine's routing status before any long-form job starts.
Streaming API Headers
Both REST and WebSocket endpoints attach routing metadata:
- OpenAI-compatible REST (
/v1/audio/speech): HeadersX-VoiceStudio-RoutingandX-VoiceStudio-Routing-Reason - WebSocket (
/ws/tts): Initial{type: "routing", status, reason}frame
# Python client handling routing headers
import httpx
def synth_longform(payload):
with httpx.Client(base_url="https://api.voicestudio.dev") as client:
r = client.post("/longform/render", json=payload, timeout=300)
routing = r.headers.get("X-VoiceStudio-Routing")
if routing == "cpu_fallback":
print("⚠️ Warning:", r.headers.get("X-VoiceStudio-Routing-Reason"))
Frontend GPU Compatibility Matrix UI
The EngineCompatibilityMatrix.jsx component renders the matrix in Settings → Engines, highlighting the effective device and displaying status badges.
Key Rendering Logic
// frontend/src/components/EngineCompatibilityMatrix.jsx (lines 28-55)
const CHIP_EFFECTIVE = 'is-effective';
const ROUTING_BADGE = {
accelerated: { tone: 'success', labelKey: 'engines.routingAccelerated' },
cpu_fallback: { tone: 'warn', labelKey: 'engines.routingCpuFallback' },
cpu_only: { tone: 'neutral', labelKey: 'engines.routingCpuOnly' },
'n/a': { tone: 'neutral', labelKey: 'engines.routingRemote' },
};
function normalizeEntry(engineData) {
const {
effective_device,
routing_status,
routing_reason
} = engineData;
return {
effectiveDevice: effective_device,
status: routing_status,
reason: routing_reason,
badge: ROUTING_BADGE[routing_status]
};
}
The component applies CHIP_EFFECTIVE to the GPU chip matching effective_device, ensuring users instantly see which accelerator is active.
Client-Side Routing Checks Before Long-Form Jobs
Implement proactive compatibility checks to warn users before starting audiobook or EPUB/PDF conversion:
// Client-side check before launching long-form generation
async function canRunLongform(engineId) {
const resp = await fetch('/engines');
const data = await resp.json();
const engine = data.tts.backends.find(e => e.id === engineId);
if (!engine) throw new Error('Engine not found');
switch (engine.routing_status) {
case 'accelerated':
return true; // Optimal path
case 'cpu_fallback':
console.warn(engine.routing_reason);
return confirm('GPU unavailable. Run on CPU (significantly slower)?');
case 'unavailable':
alert(`Cannot run on this host: ${engine.routing_reason}`);
return false;
default:
return false;
}
}
Why GPU Compatibility Matters for Long-Form Features
Long-form rendering (/longform/render, /audiobook) streams audio chunks through the GPU-pool worker. Silent CPU fallbacks cause 10× performance degradation on hour-long audiobooks or multi-chapter EPUB conversions.
The VoiceStudio GPU compatibility matrix prevents this by:
- Explicit warnings:
cpu_fallbacktriggers UI badges and pre-flight warnings - Hard blocking:
unavailablereturns HTTP 400 or WebSocket errors rather than failing mid-job - Performance transparency: Device name and VRAM display set latency expectations
Summary
- Device probe (
backend/core/device_caps.py): Detects CUDA, ROCm, MPS, XPU, or CPU with driver/VRAM diagnostics - Engine declarations (
backend/services/tts_backend.py, engine modules): Each backend lists compatible GPU families ingpu_compat - Routing resolver (
backend/services/engine_routing.py): Pure function produceseffective_device,routing_status, androuting_reason - Five routing statuses:
accelerated,cpu_fallback,cpu_only,unavailable,n/afor LLMs - Multi-interface exposure: REST
/engines, wizard pre-flight, HTTP headers, WebSocket frames, and React UI matrix - Long-form protection: Eliminates silent CPU fallbacks that would degrade audiobook and EPUB/PDF import performance by 10×
Frequently Asked Questions
What GPU families does VoiceStudio support for long-form features?
VoiceStudio supports CUDA (NVIDIA), ROCm (AMD), MPS (Apple Silicon), XPU (Intel), and CPU fallback. The device probe in backend/core/device_caps.py detects all five families, with family indicating the best available accelerator and available_families listing all usable options.
How do I check if my audiobook job will run on GPU before starting?
Query the /engines endpoint and inspect routing_status for your selected TTS backend. Status accelerated confirms GPU execution. Status cpu_fallback warns of CPU degradation. The pre-flight wizard (backend/wizard.py) and streaming API headers provide the same information without manual API calls.
Why does VoiceStudio block some engines instead of falling back to CPU?
Engines with routing_status: unavailable lack both the host's GPU family and cpu in their gpu_compat declaration. VoiceStudio blocks these rather than fail mid-job. This occurs when specialized engines require specific hardware (e.g., CUDA-only models on Apple Silicon hosts without CPU support declared).
Can I extend GPU compatibility to custom TTS backends?
Yes. Add a gpu_compat tuple to your engine's initialization file (pattern: engines/your_engine/__init__.py). Include "cuda", "rocm", "mps", "xpu", and/or "cpu" based on your model's tested platforms. The routing resolver automatically incorporates your engine into the compatibility matrix without frontend changes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →