# Supported TTS Engines in VoiceStudio: A Complete Guide to 11 Backends

> Explore VoiceStudio's 11 supported TTS engines like OmniVoice, CosyVoice, and GPT-SoVITS. This guide details each backend for seamless integration into your projects.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: getting-started
- Published: 2026-09-09

---

**VoiceStudio supports 11 Text-to-Speech engines** including OmniVoice, CosyVoice, GPT-SoVITS, and MLX-Audio, each implemented as a modular backend subclassing the abstract `TTSBackend` class defined in the core service layer.

VoiceStudio is an open-source TTS orchestration platform that provides a unified interface to multiple synthesis engines. The supported TTS engines in VoiceStudio range from lightweight models optimized for edge devices to state-of-the-art neural architectures, all integrated through a dynamic backend registry system that auto-discovers available implementations at runtime.

## Complete List of Supported TTS Engines

VoiceStudio ships with 11 distinct TTS backends, each residing in the `backend/engines/` directory and subclassing the abstract `TTSBackend` class. The `BackendRegistry` automatically detects these implementations when their modules are imported.

- **OmniVoice** (`OmniVoiceBackend`): A general-purpose TTS engine based on the OmniVoice project, implemented in [`backend/engines/omnivoice_gguf/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/omnivoice_gguf/backend.py).

- **VoxCPM2** (`VoxCPM2Backend`): High-quality neural TTS utilizing the Vox-CPM-2 model, defined in [`backend/engines/voxcp2/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/voxcp2/backend.py).

- **MOSS-TTS-Nano** (`MossTTSNanoBackend`): A lightweight model optimized for low-resource devices, implemented in [`backend/engines/moss_tts_nano/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/moss_tts_nano/backend.py).

- **Kitten-TTS** (`KittenTTSBackend`): A simple reference implementation primarily used for testing and validation.

- **MLX-Audio** (`MLXAudioBackend`): Apple Silicon-optimized backend leveraging the MLX-Audio library for accelerated inference on macOS devices, found in [`backend/engines/mlx_audio/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/mlx_audio/backend.py).

- **CosyVoice** (`CosyVoiceBackend`): State-of-the-art TTS from the CosyVoice project, located in [`backend/engines/cosyvoice/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/cosyvoice/backend.py).

- **GPT-SoVITS** (`GPTSoVITSBackend`): Text-to-speech synthesis built on the GPT-SoVITS framework, implemented in [`backend/engines/gpt_sovits/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/gpt_sovits/backend.py).

- **Sherpa-Onnx** (`SherpaOnnxBackend`): On-device TTS engine using the Sherpa-Onnx runtime for offline inference, located in [`backend/engines/sherpa_onnx/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/sherpa_onnx/backend.py).

- **Pocket-TTS** (`PocketTTSBackend`): A small, fast subprocess-based engine defined in [`backend/engines/pockettts/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/pockettts/__init__.py).

- **Dots-TTS** (`DotsTTSBackend`): A minimal backend for quick demonstrations and prototyping, implemented in [`backend/engines/dots_tts/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/dots_tts/__init__.py).

- **AudioCPP** (`AudioCPPBackend`): C++-based TTS wrapped for Python usage, located in [`backend/engines/audiocpp/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/audiocpp/__init__.py).

## Architecture and Backend Registration

All engines inherit from the abstract base class `TTSBackend` defined in [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py) at line 205. This class establishes the contract that every TTS engine must implement, ensuring consistent method signatures across the ecosystem.

The `BackendRegistry` discovers concrete implementations automatically at import time. Each backend module registers itself upon initialization, making the engine available for selection via the `OMNIVOICE_TTS_BACKEND` environment variable. This variable determines which backend class is instantiated when `get_backend()` is called.

```python

# backend/services/tts_backend.py

class TTSBackend(ABC):
    @abstractmethod
    def synthesize(self, text: str) -> bytes:
        ...

```

## Selecting and Using TTS Engines

You control the active TTS engine by setting the `OMNIVOICE_TTS_BACKEND` environment variable before initializing the backend. The registry maps this string to the corresponding backend class.

**Selecting an engine:**

```python
import os

# Select CosyVoice as the active backend

os.environ["OMNIVOICE_TTS_BACKEND"] = "cosyvoice"

```

**Generating speech:**

```python
from backend.services.tts_backend import get_backend

backend = get_backend()  # Returns the active TTSBackend instance

audio = backend.synthesize("Hello, world!")

```

**Listing available engines:**

```python
from backend.services.tts_backend import BackendRegistry

print("Available TTS engines:")
for name, cls in BackendRegistry.backends().items():
    print(f"- {name}: {cls.__name__}")

```

## Engine Categories and Use Cases

### High-Quality Neural Synthesis

**CosyVoice** and **VoxCPM2** provide state-of-the-art speech quality for production applications requiring natural prosody. **GPT-SoVITS** offers zero-shot voice cloning capabilities through its GPT-based architecture.

### Edge and Mobile Deployment

**MOSS-TTS-Nano**, **Pocket-TTS**, and **Sherpa-Onnx** are optimized for resource-constrained environments. Sherpa-Onnx runs entirely on-device without cloud dependencies, while Pocket-TTS operates as a lightweight subprocess ideal for embedded systems.

### Platform-Specific Optimization

**MLX-Audio** specifically targets Apple Silicon (M1/M2/M3) processors, utilizing the MLX framework to achieve hardware-accelerated inference on macOS devices.

### Development and Testing

**Kitten-TTS** and **Dots-TTS** serve as reference implementations. These minimal backends facilitate rapid prototyping and unit testing without requiring large model downloads or GPU resources.

## Summary

- VoiceStudio supports **11 distinct TTS engines** through a modular backend architecture.
- All engines implement the abstract `TTSBackend` class defined in [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py).
- The `BackendRegistry` auto-discovers backends; selection occurs via the `OMNIVOICE_TTS_BACKEND` environment variable.
- Engines span categories from lightweight edge models (**MOSS-TTS-Nano**, **Pocket-TTS**) to high-fidelity neural synthesis (**CosyVoice**, **GPT-SoVITS**).
- Platform-specific optimizations include **MLX-Audio** for Apple Silicon and **Sherpa-Onnx** for offline on-device inference.
- Adding new engines requires subclassing `TTSBackend` and ensuring the module is importable by the registry.

## Frequently Asked Questions

### How do I add a custom TTS engine to VoiceStudio?

Subclass the abstract `TTSBackend` class from [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py) and implement the required `synthesize` method. Place your module in the `backend/engines/` directory or ensure it is importable at runtime. The `BackendRegistry` automatically discovers and registers any imported subclass of `TTSBackend`.

### Which TTS engine should I use for Apple Silicon Macs?

Use **MLX-Audio** (`MLXAudioBackend`) for optimal performance on Apple Silicon. This backend leverages the MLX framework for hardware-accelerated inference on M1, M2, and M3 processors, providing significantly faster synthesis compared to CPU-bound alternatives.

### Can I switch TTS engines at runtime?

Switching engines requires setting the `OMNIVOICE_TTS_BACKEND` environment variable and obtaining a new backend instance via `get_backend()`. While the registry supports dynamic backend discovery, the active synthesis instance is determined at initialization time, so you must reinitialize the backend to switch engines.

### Where are the TTS engine implementations stored?

Each engine resides in its own subdirectory under `backend/engines/`. For example, Pocket-TTS is defined in [`backend/engines/pockettts/__init__.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/pockettts/__init__.py), while CosyVoice is implemented in [`backend/engines/cosyvoice/backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/engines/cosyvoice/backend.py). The abstract base class that all engines extend is located at [`backend/services/tts_backend.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/services/tts_backend.py).