Supported TTS Engines in VoiceStudio: A Complete Guide to 11 Backends
VoiceStudio supports 11 Text-to-Speech engines including OmniVoice, CosyVoice, GPT-SoVITS, and MLX-Audio, each implemented as a modular backend subclassing the abstract TTSBackend class defined in the core service layer.
VoiceStudio is an open-source TTS orchestration platform that provides a unified interface to multiple synthesis engines. The supported TTS engines in VoiceStudio range from lightweight models optimized for edge devices to state-of-the-art neural architectures, all integrated through a dynamic backend registry system that auto-discovers available implementations at runtime.
Complete List of Supported TTS Engines
VoiceStudio ships with 11 distinct TTS backends, each residing in the backend/engines/ directory and subclassing the abstract TTSBackend class. The BackendRegistry automatically detects these implementations when their modules are imported.
-
OmniVoice (
OmniVoiceBackend): A general-purpose TTS engine based on the OmniVoice project, implemented inbackend/engines/omnivoice_gguf/backend.py. -
VoxCPM2 (
VoxCPM2Backend): High-quality neural TTS utilizing the Vox-CPM-2 model, defined inbackend/engines/voxcp2/backend.py. -
MOSS-TTS-Nano (
MossTTSNanoBackend): A lightweight model optimized for low-resource devices, implemented inbackend/engines/moss_tts_nano/backend.py. -
Kitten-TTS (
KittenTTSBackend): A simple reference implementation primarily used for testing and validation. -
MLX-Audio (
MLXAudioBackend): Apple Silicon-optimized backend leveraging the MLX-Audio library for accelerated inference on macOS devices, found inbackend/engines/mlx_audio/backend.py. -
CosyVoice (
CosyVoiceBackend): State-of-the-art TTS from the CosyVoice project, located inbackend/engines/cosyvoice/backend.py. -
GPT-SoVITS (
GPTSoVITSBackend): Text-to-speech synthesis built on the GPT-SoVITS framework, implemented inbackend/engines/gpt_sovits/backend.py. -
Sherpa-Onnx (
SherpaOnnxBackend): On-device TTS engine using the Sherpa-Onnx runtime for offline inference, located inbackend/engines/sherpa_onnx/backend.py. -
Pocket-TTS (
PocketTTSBackend): A small, fast subprocess-based engine defined inbackend/engines/pockettts/__init__.py. -
Dots-TTS (
DotsTTSBackend): A minimal backend for quick demonstrations and prototyping, implemented inbackend/engines/dots_tts/__init__.py. -
AudioCPP (
AudioCPPBackend): C++-based TTS wrapped for Python usage, located inbackend/engines/audiocpp/__init__.py.
Architecture and Backend Registration
All engines inherit from the abstract base class TTSBackend defined in backend/services/tts_backend.py at line 205. This class establishes the contract that every TTS engine must implement, ensuring consistent method signatures across the ecosystem.
The BackendRegistry discovers concrete implementations automatically at import time. Each backend module registers itself upon initialization, making the engine available for selection via the OMNIVOICE_TTS_BACKEND environment variable. This variable determines which backend class is instantiated when get_backend() is called.
# backend/services/tts_backend.py
class TTSBackend(ABC):
@abstractmethod
def synthesize(self, text: str) -> bytes:
...
Selecting and Using TTS Engines
You control the active TTS engine by setting the OMNIVOICE_TTS_BACKEND environment variable before initializing the backend. The registry maps this string to the corresponding backend class.
Selecting an engine:
import os
# Select CosyVoice as the active backend
os.environ["OMNIVOICE_TTS_BACKEND"] = "cosyvoice"
Generating speech:
from backend.services.tts_backend import get_backend
backend = get_backend() # Returns the active TTSBackend instance
audio = backend.synthesize("Hello, world!")
Listing available engines:
from backend.services.tts_backend import BackendRegistry
print("Available TTS engines:")
for name, cls in BackendRegistry.backends().items():
print(f"- {name}: {cls.__name__}")
Engine Categories and Use Cases
High-Quality Neural Synthesis
CosyVoice and VoxCPM2 provide state-of-the-art speech quality for production applications requiring natural prosody. GPT-SoVITS offers zero-shot voice cloning capabilities through its GPT-based architecture.
Edge and Mobile Deployment
MOSS-TTS-Nano, Pocket-TTS, and Sherpa-Onnx are optimized for resource-constrained environments. Sherpa-Onnx runs entirely on-device without cloud dependencies, while Pocket-TTS operates as a lightweight subprocess ideal for embedded systems.
Platform-Specific Optimization
MLX-Audio specifically targets Apple Silicon (M1/M2/M3) processors, utilizing the MLX framework to achieve hardware-accelerated inference on macOS devices.
Development and Testing
Kitten-TTS and Dots-TTS serve as reference implementations. These minimal backends facilitate rapid prototyping and unit testing without requiring large model downloads or GPU resources.
Summary
- VoiceStudio supports 11 distinct TTS engines through a modular backend architecture.
- All engines implement the abstract
TTSBackendclass defined inbackend/services/tts_backend.py. - The
BackendRegistryauto-discovers backends; selection occurs via theOMNIVOICE_TTS_BACKENDenvironment variable. - Engines span categories from lightweight edge models (MOSS-TTS-Nano, Pocket-TTS) to high-fidelity neural synthesis (CosyVoice, GPT-SoVITS).
- Platform-specific optimizations include MLX-Audio for Apple Silicon and Sherpa-Onnx for offline on-device inference.
- Adding new engines requires subclassing
TTSBackendand ensuring the module is importable by the registry.
Frequently Asked Questions
How do I add a custom TTS engine to VoiceStudio?
Subclass the abstract TTSBackend class from backend/services/tts_backend.py and implement the required synthesize method. Place your module in the backend/engines/ directory or ensure it is importable at runtime. The BackendRegistry automatically discovers and registers any imported subclass of TTSBackend.
Which TTS engine should I use for Apple Silicon Macs?
Use MLX-Audio (MLXAudioBackend) for optimal performance on Apple Silicon. This backend leverages the MLX framework for hardware-accelerated inference on M1, M2, and M3 processors, providing significantly faster synthesis compared to CPU-bound alternatives.
Can I switch TTS engines at runtime?
Switching engines requires setting the OMNIVOICE_TTS_BACKEND environment variable and obtaining a new backend instance via get_backend(). While the registry supports dynamic backend discovery, the active synthesis instance is determined at initialization time, so you must reinitialize the backend to switch engines.
Where are the TTS engine implementations stored?
Each engine resides in its own subdirectory under backend/engines/. For example, Pocket-TTS is defined in backend/engines/pockettts/__init__.py, while CosyVoice is implemented in backend/engines/cosyvoice/backend.py. The abstract base class that all engines extend is located at backend/services/tts_backend.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →