# How to Add a New TTS Engine to Voicebox: A Complete Integration Guide

> Learn how to add a new TTS engine to Voicebox. This guide covers implementing TTSBackend protocol, registry updates, TypeScript types, UI, and PyInstaller bundling. Integrate seamlessly.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: how-to-guide
- Published: 2026-04-14

---

**Adding a new TTS engine to Voicebox requires implementing a `TTSBackend` protocol class, registering it in the global engine registry, updating TypeScript types and UI components, and configuring PyInstaller bundling.**

Voicebox uses a **registry-driven architecture** that decouples HTTP routes from concrete text-to-speech implementations. According to the `jamiepine/voicebox` source code, the backend factory in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) dispatches generation requests via a global `TTS_ENGINES` mapping, allowing new engines to integrate without modifying route handlers. This guide walks through the six-phase integration process documented in `docs/content/docs/developer/tts-engines.mdx`, covering backend implementation, frontend wiring, and binary packaging.

## Understanding the Registry Architecture

Voicebox organizes TTS support into four layers that communicate through a centralized registry:

| Layer | Responsibility | Primary Files |
|------|----------------|---------------|
| **Routes / Services** | Thin HTTP handlers that delegate to the backend via the model-config registry. No per-engine code required. | `backend/routes/*`, `backend/services/*` |
| **Backends** | Engine-specific implementations of the `TTSBackend` protocol that register a `ModelConfig` defining model metadata. | [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py), `backend/backends/<engine>_backend.py` |
| **Frontend** | UI selectors, TypeScript types, and language maps exposing the engine to users. | [`app/src/components/Generation/EngineModelSelector.tsx`](https://github.com/jamiepine/voicebox/blob/main/app/src/components/Generation/EngineModelSelector.tsx), [`app/src/lib/api/types.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/api/types.ts) |
| **Packaging** | PyInstaller bundling, dependency auditing, and CI configuration. | [`backend/build_binary.py`](https://github.com/jamiepine/voicebox/blob/main/backend/build_binary.py), [`backend/requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements.txt), [`.github/workflows/release.yml`](https://github.com/jamiepine/voicebox/blob/main/.github/workflows/release.yml) |

The **registry** in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) defines the `ModelConfig` dataclass and the global `TTS_ENGINES` dictionary. The factory function `get_tts_backend_for_engine()` (around line 511) performs a cached singleton lookup, returning the concrete backend instance so HTTP routes only call `backend.get_tts_backend().generate(...)` without engine-specific logic.

## The Six-Phase Integration Workflow

The Voicebox documentation prescribes a phased approach to ensure engines work in both development and frozen binaries:

| Phase | Goal | Files to Modify |
|------|------|----------------|
| **0 – Dependency Research** | Audit the third-party library for PyInstaller compatibility, native data files, required monkey-patches, and download methods. | No code changes—produce a written audit. |
| **1 – Backend Implementation** | Create a class satisfying the `TTSBackend` protocol, add a `ModelConfig`, and register the engine. | `backend/backends/<engine>_backend.py`, [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) |
| **2 – Route / Service** | No changes needed—the registry auto-dispatches. | None |
| **3 – Frontend Wiring** | Extend TypeScript union types, language maps, and UI selectors. | [`app/src/lib/api/types.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/api/types.ts), [`app/src/lib/constants/languages.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/constants/languages.ts), [`app/src/components/Generation/EngineModelSelector.tsx`](https://github.com/jamiepine/voicebox/blob/main/app/src/components/Generation/EngineModelSelector.tsx) |
| **4 – Dependency Declaration** | Add Python packages to [`backend/requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements.txt), `justfile`, CI workflow, and `Dockerfile`. | [`backend/requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements.txt), `justfile`, [`.github/workflows/release.yml`](https://github.com/jamiepine/voicebox/blob/main/.github/workflows/release.yml) |
| **5 – PyInstaller Bundling** | Add hidden-imports and collect-all directives for the new engine's packages. | [`backend/build_binary.py`](https://github.com/jamiepine/voicebox/blob/main/backend/build_binary.py), [`backend/server.py`](https://github.com/jamiepine/voicebox/blob/main/backend/server.py) |
| **6 – Common Upstream Workarounds** | Apply monkey-patches discovered in Phase 0 (e.g., `torch.load` map-location fixes). | `backend/backends/<engine>_backend.py` |

## Phase 1: Implementing the TTS Backend

Create [`backend/backends/myengine_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/myengine_backend.py) implementing the `TTSBackend` protocol. The class must provide `MODEL_CONFIGS`, handle model loading, voice prompt creation, and audio generation.

```python
from __future__ import annotations
import numpy as np
from typing import List, Tuple

from . import TTSBackend, ModelConfig
from ..utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt

class MyEngineBackend:
    """Concrete implementation for the MyEngine TTS library."""
    
    MODEL_CONFIGS: List[ModelConfig] = [
        ModelConfig(
            model_name="myengine-tts-1B",
            display_name="MyEngine TTS 1B",
            engine="myengine",
            hf_repo_id="myorg/myengine-tts-1b",
            size_mb=2500,
            languages=["en", "fr", "de"],
        )
    ]

    def __init__(self) -> None:
        self._model = None

    async def load_model(self, model_size: str = "default") -> None:
        from myengine import MyEngineModel
        repo = self._model_config.hf_repo_id
        self._model = MyEngineModel.from_pretrained(repo, device="cpu")

    async def create_voice_prompt(
        self,
        audio_path: str,
        reference_text: str,
        use_cache: bool = True,
    ) -> Tuple[dict, bool]:
        cache_key = "myengine_" + get_cache_key(audio_path, reference_text)
        if use_cache:
            cached = get_cached_voice_prompt(cache_key)
            if cached:
                return cached, True
        
        prompt = self._model.encode_reference(audio_path)
        data = {"prompt": prompt, "type": "cloned"}
        cache_voice_prompt(cache_key, data)
        return data, False

    async def combine_voice_prompts(
        self,
        audio_paths: List[str],
        reference_texts: List[str],
    ) -> Tuple[np.ndarray, str]:
        prompts = [self._model.encode_reference(p) for p in audio_paths]
        combined = np.mean(np.stack(prompts), axis=0)
        combined_text = " ".join(reference_texts)
        return combined, combined_text

    async def generate(
        self,
        text: str,
        voice_prompt: dict,
        language: str = "en",
        seed: int | None = None,
        instruct: str | None = None,
    ) -> Tuple[np.ndarray, int]:
        audio, sr = self._model.generate(
            text,
            speaker_prompt=voice_prompt["prompt"],
            language=language,
            seed=seed,
            instruct=instruct,
        )
        return audio, sr

    def unload_model(self) -> None:
        self._model = None

    def is_loaded(self) -> bool:
        return self._model is not None

    @property
    def _model_config(self) -> ModelConfig:
        return self.MODEL_CONFIGS[0]

```

## Phase 1 (Continued): Registering the Engine

Modify [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) to include the new engine in three locations: the `ModelConfig` list, the `TTS_ENGINES` dictionary, and the factory dispatch chain.

```python

# 1. Add ModelConfig entry

ModelConfig(
    model_name="myengine-tts-1B",
    display_name="MyEngine TTS 1B",
    engine="myengine",
    hf_repo_id="myorg/myengine-tts-1b",
    size_mb=2500,
    languages=["en", "fr", "de"],
),

# 2. Add to TTS_ENGINES dict

TTS_ENGINES = {
    # ... existing engines ...

    "myengine": "MyEngine",
}

# 3. Add factory branch (around line 511)

elif engine == "myengine":
    from .myengine_backend import MyEngineBackend
    backend = MyEngineBackend()

```

## Phase 3: Frontend Integration

Update the React frontend to expose the new engine in the generation UI.

**Engine Model Selector** ([`app/src/components/Generation/EngineModelSelector.tsx`](https://github.com/jamiepine/voicebox/blob/main/app/src/components/Generation/EngineModelSelector.tsx)):

```typescript
// Add to ENGINE_OPTIONS
{ value: 'myengine', label: 'MyEngine', engine: 'myengine' },

// Add description
myengine: 'Fast, multilingual, 1B-param model',

```

If the engine only supports English, add it to the `ENGLISH_ONLY_ENGINES` set in the same file.

**TypeScript Types** ([`app/src/lib/api/types.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/api/types.ts)):

```typescript
export type GenerationRequest = {
  // ...
  engine:
    | 'qwen'
    | 'luxtts'
    // ...
    | 'myengine';          // ← new union member
  // ...
};

```

**Language Constants** ([`app/src/lib/constants/languages.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/constants/languages.ts)):

```typescript
export const ENGINE_LANGUAGES: Record<string, string[]> = {
  // ...
  myengine: ['en', 'fr', 'de'],
};

```

## Phases 4-5: Dependency and Packaging Configuration

**Python Requirements** ([`backend/requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/backend/requirements.txt)):

Add the new TTS library and any pinned sub-dependencies. If the library pins an incompatible torch version, use `--no-deps` in the `justfile` and list sub-dependencies manually.

**PyInstaller Configuration** ([`backend/build_binary.py`](https://github.com/jamiepine/voicebox/blob/main/backend/build_binary.py)):

Add hidden-imports and collect-all directives to include the backend module and any data files:

```python
hidden_imports.append('backend.backends.myengine_backend')
collect_all.append('myengine')               # package with data files

copy_metadata.append('myengine')             # if using importlib.metadata

```

**CI and Docker** ([`.github/workflows/release.yml`](https://github.com/jamiepine/voicebox/blob/main/.github/workflows/release.yml), `Dockerfile`):

Mirror the install commands from [`requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/requirements.txt) and `justfile` to ensure production builds match the development environment.

## Summary

- **Voicebox uses a registry pattern** where [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) maps engine keys to singleton backend instances via `get_tts_backend_for_engine()`.
- **Backend implementation** requires creating a class satisfying the `TTSBackend` protocol with `load_model()`, `generate()`, and voice prompt methods.
- **Registration** involves updating `ModelConfig`, `TTS_ENGINES`, and the factory chain in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py).
- **Frontend updates** include TypeScript union types in [`app/src/lib/api/types.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/api/types.ts), language maps in [`app/src/lib/constants/languages.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/constants/languages.ts), and UI options in [`EngineModelSelector.tsx`](https://github.com/jamiepine/voicebox/blob/main/EngineModelSelector.tsx).
- **Packaging** requires PyInstaller hidden-imports in [`backend/build_binary.py`](https://github.com/jamiepine/voicebox/blob/main/backend/build_binary.py) and dependency declarations in [`requirements.txt`](https://github.com/jamiepine/voicebox/blob/main/requirements.txt), `justfile`, and CI workflows.

## Frequently Asked Questions

### What is the TTSBackend protocol in Voicebox?

The `TTSBackend` protocol is an informal interface defined in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) that concrete engine classes must implement. It requires methods for `load_model()`, `generate()`, `create_voice_prompt()`, `combine_voice_prompts()`, `unload_model()`, and `is_loaded()`. The registry factory returns instances satisfying this protocol, allowing HTTP routes to call generation logic without knowing the specific engine implementation.

### Do I need to modify HTTP routes when adding a new engine?

No. The HTTP routes in `backend/routes/` delegate to `backend/services/`, which calls `get_tts_backend().generate()`. Because the factory in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) handles engine dispatch via the `TTS_ENGINES` registry, routes automatically support new engines once registered without code changes.

### How do I handle engines that require specific PyInstaller hooks?

Add `--hidden-import` directives for the backend module path in [`backend/build_binary.py`](https://github.com/jamiepine/voicebox/blob/main/backend/build_binary.py). If the engine's package loads data files or uses `inspect.getsource()`, add it to the `collect_all` list. For libraries using `importlib.metadata`, include them in `copy_metadata`. Test frozen builds with `just build` before releasing.

### Where does Voicebox store the supported languages for each engine?

Language support is defined in two places: the `languages` field of the `ModelConfig` dataclass in the Python backend, and the `ENGINE_LANGUAGES` record in [`app/src/lib/constants/languages.ts`](https://github.com/jamiepine/voicebox/blob/main/app/src/lib/constants/languages.ts) on the frontend. Both must be updated to ensure the UI correctly filters language options and the backend validates requests.