How to Add a New TTS Engine to Voicebox: A Complete Integration Guide
Adding a new TTS engine to Voicebox requires implementing a TTSBackend protocol class, registering it in the global engine registry, updating TypeScript types and UI components, and configuring PyInstaller bundling.
Voicebox uses a registry-driven architecture that decouples HTTP routes from concrete text-to-speech implementations. According to the jamiepine/voicebox source code, the backend factory in backend/backends/__init__.py dispatches generation requests via a global TTS_ENGINES mapping, allowing new engines to integrate without modifying route handlers. This guide walks through the six-phase integration process documented in docs/content/docs/developer/tts-engines.mdx, covering backend implementation, frontend wiring, and binary packaging.
Understanding the Registry Architecture
Voicebox organizes TTS support into four layers that communicate through a centralized registry:
| Layer | Responsibility | Primary Files |
|---|---|---|
| Routes / Services | Thin HTTP handlers that delegate to the backend via the model-config registry. No per-engine code required. | backend/routes/*, backend/services/* |
| Backends | Engine-specific implementations of the TTSBackend protocol that register a ModelConfig defining model metadata. |
backend/backends/__init__.py, backend/backends/<engine>_backend.py |
| Frontend | UI selectors, TypeScript types, and language maps exposing the engine to users. | app/src/components/Generation/EngineModelSelector.tsx, app/src/lib/api/types.ts |
| Packaging | PyInstaller bundling, dependency auditing, and CI configuration. | backend/build_binary.py, backend/requirements.txt, .github/workflows/release.yml |
The registry in backend/backends/__init__.py defines the ModelConfig dataclass and the global TTS_ENGINES dictionary. The factory function get_tts_backend_for_engine() (around line 511) performs a cached singleton lookup, returning the concrete backend instance so HTTP routes only call backend.get_tts_backend().generate(...) without engine-specific logic.
The Six-Phase Integration Workflow
The Voicebox documentation prescribes a phased approach to ensure engines work in both development and frozen binaries:
| Phase | Goal | Files to Modify |
|---|---|---|
| 0 – Dependency Research | Audit the third-party library for PyInstaller compatibility, native data files, required monkey-patches, and download methods. | No code changes—produce a written audit. |
| 1 – Backend Implementation | Create a class satisfying the TTSBackend protocol, add a ModelConfig, and register the engine. |
backend/backends/<engine>_backend.py, backend/backends/__init__.py |
| 2 – Route / Service | No changes needed—the registry auto-dispatches. | None |
| 3 – Frontend Wiring | Extend TypeScript union types, language maps, and UI selectors. | app/src/lib/api/types.ts, app/src/lib/constants/languages.ts, app/src/components/Generation/EngineModelSelector.tsx |
| 4 – Dependency Declaration | Add Python packages to backend/requirements.txt, justfile, CI workflow, and Dockerfile. |
backend/requirements.txt, justfile, .github/workflows/release.yml |
| 5 – PyInstaller Bundling | Add hidden-imports and collect-all directives for the new engine's packages. | backend/build_binary.py, backend/server.py |
| 6 – Common Upstream Workarounds | Apply monkey-patches discovered in Phase 0 (e.g., torch.load map-location fixes). |
backend/backends/<engine>_backend.py |
Phase 1: Implementing the TTS Backend
Create backend/backends/myengine_backend.py implementing the TTSBackend protocol. The class must provide MODEL_CONFIGS, handle model loading, voice prompt creation, and audio generation.
from __future__ import annotations
import numpy as np
from typing import List, Tuple
from . import TTSBackend, ModelConfig
from ..utils.cache import get_cache_key, get_cached_voice_prompt, cache_voice_prompt
class MyEngineBackend:
"""Concrete implementation for the MyEngine TTS library."""
MODEL_CONFIGS: List[ModelConfig] = [
ModelConfig(
model_name="myengine-tts-1B",
display_name="MyEngine TTS 1B",
engine="myengine",
hf_repo_id="myorg/myengine-tts-1b",
size_mb=2500,
languages=["en", "fr", "de"],
)
]
def __init__(self) -> None:
self._model = None
async def load_model(self, model_size: str = "default") -> None:
from myengine import MyEngineModel
repo = self._model_config.hf_repo_id
self._model = MyEngineModel.from_pretrained(repo, device="cpu")
async def create_voice_prompt(
self,
audio_path: str,
reference_text: str,
use_cache: bool = True,
) -> Tuple[dict, bool]:
cache_key = "myengine_" + get_cache_key(audio_path, reference_text)
if use_cache:
cached = get_cached_voice_prompt(cache_key)
if cached:
return cached, True
prompt = self._model.encode_reference(audio_path)
data = {"prompt": prompt, "type": "cloned"}
cache_voice_prompt(cache_key, data)
return data, False
async def combine_voice_prompts(
self,
audio_paths: List[str],
reference_texts: List[str],
) -> Tuple[np.ndarray, str]:
prompts = [self._model.encode_reference(p) for p in audio_paths]
combined = np.mean(np.stack(prompts), axis=0)
combined_text = " ".join(reference_texts)
return combined, combined_text
async def generate(
self,
text: str,
voice_prompt: dict,
language: str = "en",
seed: int | None = None,
instruct: str | None = None,
) -> Tuple[np.ndarray, int]:
audio, sr = self._model.generate(
text,
speaker_prompt=voice_prompt["prompt"],
language=language,
seed=seed,
instruct=instruct,
)
return audio, sr
def unload_model(self) -> None:
self._model = None
def is_loaded(self) -> bool:
return self._model is not None
@property
def _model_config(self) -> ModelConfig:
return self.MODEL_CONFIGS[0]
Phase 1 (Continued): Registering the Engine
Modify backend/backends/__init__.py to include the new engine in three locations: the ModelConfig list, the TTS_ENGINES dictionary, and the factory dispatch chain.
# 1. Add ModelConfig entry
ModelConfig(
model_name="myengine-tts-1B",
display_name="MyEngine TTS 1B",
engine="myengine",
hf_repo_id="myorg/myengine-tts-1b",
size_mb=2500,
languages=["en", "fr", "de"],
),
# 2. Add to TTS_ENGINES dict
TTS_ENGINES = {
# ... existing engines ...
"myengine": "MyEngine",
}
# 3. Add factory branch (around line 511)
elif engine == "myengine":
from .myengine_backend import MyEngineBackend
backend = MyEngineBackend()
Phase 3: Frontend Integration
Update the React frontend to expose the new engine in the generation UI.
Engine Model Selector (app/src/components/Generation/EngineModelSelector.tsx):
// Add to ENGINE_OPTIONS
{ value: 'myengine', label: 'MyEngine', engine: 'myengine' },
// Add description
myengine: 'Fast, multilingual, 1B-param model',
If the engine only supports English, add it to the ENGLISH_ONLY_ENGINES set in the same file.
TypeScript Types (app/src/lib/api/types.ts):
export type GenerationRequest = {
// ...
engine:
| 'qwen'
| 'luxtts'
// ...
| 'myengine'; // ← new union member
// ...
};
Language Constants (app/src/lib/constants/languages.ts):
export const ENGINE_LANGUAGES: Record<string, string[]> = {
// ...
myengine: ['en', 'fr', 'de'],
};
Phases 4-5: Dependency and Packaging Configuration
Python Requirements (backend/requirements.txt):
Add the new TTS library and any pinned sub-dependencies. If the library pins an incompatible torch version, use --no-deps in the justfile and list sub-dependencies manually.
PyInstaller Configuration (backend/build_binary.py):
Add hidden-imports and collect-all directives to include the backend module and any data files:
hidden_imports.append('backend.backends.myengine_backend')
collect_all.append('myengine') # package with data files
copy_metadata.append('myengine') # if using importlib.metadata
CI and Docker (.github/workflows/release.yml, Dockerfile):
Mirror the install commands from requirements.txt and justfile to ensure production builds match the development environment.
Summary
- Voicebox uses a registry pattern where
backend/backends/__init__.pymaps engine keys to singleton backend instances viaget_tts_backend_for_engine(). - Backend implementation requires creating a class satisfying the
TTSBackendprotocol withload_model(),generate(), and voice prompt methods. - Registration involves updating
ModelConfig,TTS_ENGINES, and the factory chain inbackend/backends/__init__.py. - Frontend updates include TypeScript union types in
app/src/lib/api/types.ts, language maps inapp/src/lib/constants/languages.ts, and UI options inEngineModelSelector.tsx. - Packaging requires PyInstaller hidden-imports in
backend/build_binary.pyand dependency declarations inrequirements.txt,justfile, and CI workflows.
Frequently Asked Questions
What is the TTSBackend protocol in Voicebox?
The TTSBackend protocol is an informal interface defined in backend/backends/__init__.py that concrete engine classes must implement. It requires methods for load_model(), generate(), create_voice_prompt(), combine_voice_prompts(), unload_model(), and is_loaded(). The registry factory returns instances satisfying this protocol, allowing HTTP routes to call generation logic without knowing the specific engine implementation.
Do I need to modify HTTP routes when adding a new engine?
No. The HTTP routes in backend/routes/ delegate to backend/services/, which calls get_tts_backend().generate(). Because the factory in backend/backends/__init__.py handles engine dispatch via the TTS_ENGINES registry, routes automatically support new engines once registered without code changes.
How do I handle engines that require specific PyInstaller hooks?
Add --hidden-import directives for the backend module path in backend/build_binary.py. If the engine's package loads data files or uses inspect.getsource(), add it to the collect_all list. For libraries using importlib.metadata, include them in copy_metadata. Test frozen builds with just build before releasing.
Where does Voicebox store the supported languages for each engine?
Language support is defined in two places: the languages field of the ModelConfig dataclass in the Python backend, and the ENGINE_LANGUAGES record in app/src/lib/constants/languages.ts on the frontend. Both must be updated to ensure the UI correctly filters language options and the backend validates requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →