Supported TTS Engines in VoiceStudio: A Complete Guide to 11 Backends

VoiceStudio supports 11 Text-to-Speech engines including OmniVoice, CosyVoice, GPT-SoVITS, and MLX-Audio, each implemented as a modular backend subclassing the abstract TTSBackend class defined in the core service layer.

VoiceStudio is an open-source TTS orchestration platform that provides a unified interface to multiple synthesis engines. The supported TTS engines in VoiceStudio range from lightweight models optimized for edge devices to state-of-the-art neural architectures, all integrated through a dynamic backend registry system that auto-discovers available implementations at runtime.

Complete List of Supported TTS Engines

VoiceStudio ships with 11 distinct TTS backends, each residing in the backend/engines/ directory and subclassing the abstract TTSBackend class. The BackendRegistry automatically detects these implementations when their modules are imported.

Architecture and Backend Registration

All engines inherit from the abstract base class TTSBackend defined in backend/services/tts_backend.py at line 205. This class establishes the contract that every TTS engine must implement, ensuring consistent method signatures across the ecosystem.

The BackendRegistry discovers concrete implementations automatically at import time. Each backend module registers itself upon initialization, making the engine available for selection via the OMNIVOICE_TTS_BACKEND environment variable. This variable determines which backend class is instantiated when get_backend() is called.


# backend/services/tts_backend.py

class TTSBackend(ABC):
    @abstractmethod
    def synthesize(self, text: str) -> bytes:
        ...

Selecting and Using TTS Engines

You control the active TTS engine by setting the OMNIVOICE_TTS_BACKEND environment variable before initializing the backend. The registry maps this string to the corresponding backend class.

Selecting an engine:

import os

# Select CosyVoice as the active backend

os.environ["OMNIVOICE_TTS_BACKEND"] = "cosyvoice"

Generating speech:

from backend.services.tts_backend import get_backend

backend = get_backend()  # Returns the active TTSBackend instance

audio = backend.synthesize("Hello, world!")

Listing available engines:

from backend.services.tts_backend import BackendRegistry

print("Available TTS engines:")
for name, cls in BackendRegistry.backends().items():
    print(f"- {name}: {cls.__name__}")

Engine Categories and Use Cases

High-Quality Neural Synthesis

CosyVoice and VoxCPM2 provide state-of-the-art speech quality for production applications requiring natural prosody. GPT-SoVITS offers zero-shot voice cloning capabilities through its GPT-based architecture.

Edge and Mobile Deployment

MOSS-TTS-Nano, Pocket-TTS, and Sherpa-Onnx are optimized for resource-constrained environments. Sherpa-Onnx runs entirely on-device without cloud dependencies, while Pocket-TTS operates as a lightweight subprocess ideal for embedded systems.

Platform-Specific Optimization

MLX-Audio specifically targets Apple Silicon (M1/M2/M3) processors, utilizing the MLX framework to achieve hardware-accelerated inference on macOS devices.

Development and Testing

Kitten-TTS and Dots-TTS serve as reference implementations. These minimal backends facilitate rapid prototyping and unit testing without requiring large model downloads or GPU resources.

Summary

  • VoiceStudio supports 11 distinct TTS engines through a modular backend architecture.
  • All engines implement the abstract TTSBackend class defined in backend/services/tts_backend.py.
  • The BackendRegistry auto-discovers backends; selection occurs via the OMNIVOICE_TTS_BACKEND environment variable.
  • Engines span categories from lightweight edge models (MOSS-TTS-Nano, Pocket-TTS) to high-fidelity neural synthesis (CosyVoice, GPT-SoVITS).
  • Platform-specific optimizations include MLX-Audio for Apple Silicon and Sherpa-Onnx for offline on-device inference.
  • Adding new engines requires subclassing TTSBackend and ensuring the module is importable by the registry.

Frequently Asked Questions

How do I add a custom TTS engine to VoiceStudio?

Subclass the abstract TTSBackend class from backend/services/tts_backend.py and implement the required synthesize method. Place your module in the backend/engines/ directory or ensure it is importable at runtime. The BackendRegistry automatically discovers and registers any imported subclass of TTSBackend.

Which TTS engine should I use for Apple Silicon Macs?

Use MLX-Audio (MLXAudioBackend) for optimal performance on Apple Silicon. This backend leverages the MLX framework for hardware-accelerated inference on M1, M2, and M3 processors, providing significantly faster synthesis compared to CPU-bound alternatives.

Can I switch TTS engines at runtime?

Switching engines requires setting the OMNIVOICE_TTS_BACKEND environment variable and obtaining a new backend instance via get_backend(). While the registry supports dynamic backend discovery, the active synthesis instance is determined at initialization time, so you must reinitialize the backend to switch engines.

Where are the TTS engine implementations stored?

Each engine resides in its own subdirectory under backend/engines/. For example, Pocket-TTS is defined in backend/engines/pockettts/__init__.py, while CosyVoice is implemented in backend/engines/cosyvoice/backend.py. The abstract base class that all engines extend is located at backend/services/tts_backend.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →