How to Add a New TTS Engine to Voicebox: A Complete Integration Guide

Adding a new TTS engine to Voicebox requires implementing a TTSBackend protocol class, registering it in the global engine registry, updating TypeScript types and UI components, and configuring PyInstaller bundling.

Voicebox uses a registry-driven architecture that decouples HTTP routes from concrete text-to-speech implementations. According to the jamiepine/voicebox source code, the backend factory in backend/backends/__init__.py dispatches generation requests via a global TTS_ENGINES mapping, allowing new engines to integrate without modifying route handlers. This guide walks through the six-phase integration process documented in docs/content/docs/developer/tts-engines.mdx, covering backend implementation, frontend wiring, and binary packaging.

Understanding the Registry Architecture

Voicebox organizes TTS support into four layers that communicate through a centralized registry:

Layer Responsibility Primary Files
Routes / Services Thin HTTP handlers that delegate to the backend via the model-config registry. No per-engine code required. backend/routes/*, backend/services/*
Backends Engine-specific implementations of the TTSBackend protocol that register a ModelConfig defining model metadata. backend/backends/__init__.py, backend/backends/<engine>_backend.py
Frontend UI selectors, TypeScript types, and language maps exposing the engine to users. app/src/components/Generation/EngineModelSelector.tsx, app/src/lib/api/types.ts
Packaging PyInstaller bundling, dependency auditing, and CI configuration. backend/build_binary.py, backend/requirements.txt, .github/workflows/release.yml

The registry in backend/backends/__init__.py defines the ModelConfig dataclass and the global TTS_ENGINES dictionary. The factory function get_tts_backend_for_engine() (around line 511) performs a cached singleton lookup, returning the concrete backend instance so HTTP routes only call backend.get_tts_backend().generate(...) without engine-specific logic.

The Six-Phase Integration Workflow

The Voicebox documentation prescribes a phased approach to ensure engines work in both development and frozen binaries:

Phase Goal Files to Modify
0 – Dependency Research Audit the third-party library for PyInstaller compatibility, native data files, required monkey-patches, and download methods. No code changes—produce a written audit.
1 – Backend Implementation Create a class satisfying the TTSBackend protocol, add a ModelConfig, and register the engine. backend/backends/<engine>_backend.py, backend/backends/__init__.py
2 – Route / Service No changes needed—the registry auto-dispatches. None
3 – Frontend Wiring Extend TypeScript union types, language maps, and UI selectors. app/src/lib/api/types.ts, app/src/lib/constants/languages.ts, app/src/components/Generation/EngineModelSelector.tsx
4 – Dependency Declaration Add Python packages to backend/requirements.txt, justfile, CI workflow, and Dockerfile. backend/requirements.txt, justfile, .github/workflows/release.yml
5 – PyInstaller Bundling Add hidden-imports and collect-all directives for the new engine's packages. backend/build_binary.py, backend/server.py
6 – Common Upstream Workarounds Apply monkey-patches discovered in Phase 0 (e.g., torch.load map-location fixes). backend/backends/<engine>_backend.py

Phase 1: Implementing the TTS Backend

Create backend/backends/myengine_backend.py implementing the TTSBackend protocol. The class must provide MODEL_CONFIGS, handle model loading, voice prompt creation, and audio generation.

from __future__ import annotations
import numpy as np
from typing import List, Tuple

from . import TTSBackend, ModelConfig
from ..utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt

class MyEngineBackend:
    """Concrete implementation for the MyEngine TTS library."""
    
    MODEL_CONFIGS: List[ModelConfig] = [
        ModelConfig(
            model_name="myengine-tts-1B",
            display_name="MyEngine TTS 1B",
            engine="myengine",
            hf_repo_id="myorg/myengine-tts-1b",
            size_mb=2500,
            languages=["en", "fr", "de"],
        )
    ]

    def __init__(self) -> None:
        self._model = None

    async def load_model(self, model_size: str = "default") -> None:
        from myengine import MyEngineModel
        repo = self._model_config.hf_repo_id
        self._model = MyEngineModel.from_pretrained(repo, device="cpu")

    async def create_voice_prompt(
        self,
        audio_path: str,
        reference_text: str,
        use_cache: bool = True,
    ) -> Tuple[dict, bool]:
        cache_key = "myengine_" + get_cache_key(audio_path, reference_text)
        if use_cache:
            cached = get_cached_voice_prompt(cache_key)
            if cached:
                return cached, True
        
        prompt = self._model.encode_reference(audio_path)
        data = {"prompt": prompt, "type": "cloned"}
        cache_voice_prompt(cache_key, data)
        return data, False

    async def combine_voice_prompts(
        self,
        audio_paths: List[str],
        reference_texts: List[str],
    ) -> Tuple[np.ndarray, str]:
        prompts = [self._model.encode_reference(p) for p in audio_paths]
        combined = np.mean(np.stack(prompts), axis=0)
        combined_text = " ".join(reference_texts)
        return combined, combined_text

    async def generate(
        self,
        text: str,
        voice_prompt: dict,
        language: str = "en",
        seed: int | None = None,
        instruct: str | None = None,
    ) -> Tuple[np.ndarray, int]:
        audio, sr = self._model.generate(
            text,
            speaker_prompt=voice_prompt["prompt"],
            language=language,
            seed=seed,
            instruct=instruct,
        )
        return audio, sr

    def unload_model(self) -> None:
        self._model = None

    def is_loaded(self) -> bool:
        return self._model is not None

    @property
    def _model_config(self) -> ModelConfig:
        return self.MODEL_CONFIGS[0]

Phase 1 (Continued): Registering the Engine

Modify backend/backends/__init__.py to include the new engine in three locations: the ModelConfig list, the TTS_ENGINES dictionary, and the factory dispatch chain.


# 1. Add ModelConfig entry

ModelConfig(
    model_name="myengine-tts-1B",
    display_name="MyEngine TTS 1B",
    engine="myengine",
    hf_repo_id="myorg/myengine-tts-1b",
    size_mb=2500,
    languages=["en", "fr", "de"],
),

# 2. Add to TTS_ENGINES dict

TTS_ENGINES = {
    # ... existing engines ...

    "myengine": "MyEngine",
}

# 3. Add factory branch (around line 511)

elif engine == "myengine":
    from .myengine_backend import MyEngineBackend
    backend = MyEngineBackend()

Phase 3: Frontend Integration

Update the React frontend to expose the new engine in the generation UI.

Engine Model Selector (app/src/components/Generation/EngineModelSelector.tsx):

// Add to ENGINE_OPTIONS
{ value: 'myengine', label: 'MyEngine', engine: 'myengine' },

// Add description
myengine: 'Fast, multilingual, 1B-param model',

If the engine only supports English, add it to the ENGLISH_ONLY_ENGINES set in the same file.

TypeScript Types (app/src/lib/api/types.ts):

export type GenerationRequest = {
  // ...
  engine:
    | 'qwen'
    | 'luxtts'
    // ...
    | 'myengine';          // ← new union member
  // ...
};

Language Constants (app/src/lib/constants/languages.ts):

export const ENGINE_LANGUAGES: Record<string, string[]> = {
  // ...
  myengine: ['en', 'fr', 'de'],
};

Phases 4-5: Dependency and Packaging Configuration

Python Requirements (backend/requirements.txt):

Add the new TTS library and any pinned sub-dependencies. If the library pins an incompatible torch version, use --no-deps in the justfile and list sub-dependencies manually.

PyInstaller Configuration (backend/build_binary.py):

Add hidden-imports and collect-all directives to include the backend module and any data files:

hidden_imports.append('backend.backends.myengine_backend')
collect_all.append('myengine')               # package with data files

copy_metadata.append('myengine')             # if using importlib.metadata

CI and Docker (.github/workflows/release.yml, Dockerfile):

Mirror the install commands from requirements.txt and justfile to ensure production builds match the development environment.

Summary

Frequently Asked Questions

What is the TTSBackend protocol in Voicebox?

The TTSBackend protocol is an informal interface defined in backend/backends/__init__.py that concrete engine classes must implement. It requires methods for load_model(), generate(), create_voice_prompt(), combine_voice_prompts(), unload_model(), and is_loaded(). The registry factory returns instances satisfying this protocol, allowing HTTP routes to call generation logic without knowing the specific engine implementation.

Do I need to modify HTTP routes when adding a new engine?

No. The HTTP routes in backend/routes/ delegate to backend/services/, which calls get_tts_backend().generate(). Because the factory in backend/backends/__init__.py handles engine dispatch via the TTS_ENGINES registry, routes automatically support new engines once registered without code changes.

How do I handle engines that require specific PyInstaller hooks?

Add --hidden-import directives for the backend module path in backend/build_binary.py. If the engine's package loads data files or uses inspect.getsource(), add it to the collect_all list. For libraries using importlib.metadata, include them in copy_metadata. Test frozen builds with just build before releasing.

Where does Voicebox store the supported languages for each engine?

Language support is defined in two places: the languages field of the ModelConfig dataclass in the Python backend, and the ENGINE_LANGUAGES record in app/src/lib/constants/languages.ts on the frontend. Both must be updated to ensure the UI correctly filters language options and the backend validates requests.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →