VoiceStudio Plugin SDK Contract for Shipping New TTS Engines: A Complete Implementation Guide
To ship a new TTS engine in VoiceStudio, implement the abstract base class TTSPlugin from backend/services/plugin_sdk.py, define required class attributes, implement three abstract methods (is_available, generate, list_voices), and decorate your class with @register_plugin in a file under backend/plugins/.
VoiceStudio is an open-source text-to-speech platform that exposes a plugin architecture through its Python backend. The VoiceStudio plugin SDK contract defines the exact interface developers must implement to add new TTS engines without modifying core application code. By following the abstract base class defined in backend/services/plugin_sdk.py, contributors can register custom synthesis engines that automatically appear in the frontend UI and integrate seamlessly with existing API endpoints.
Core Requirements of the TTSPlugin Contract
The Abstract Base Class
The foundation of the contract is the TTSPlugin abstract base class located in backend/services/plugin_sdk.py at line 73. Every new engine must create a concrete subclass that inherits from TTSPlugin and implements all abstract methods to satisfy the interface requirements.
Required Class Attributes
Your plugin class must define five specific class attributes that the SDK uses for registration and UI rendering:
id: A unique, lowercase string identifier used in API calls (line 80).display_name: The human-readable label shown in the frontend settings panel (line 84).requires_api_key: Boolean indicating whether the engine needs external API credentials (line 86).is_local: Boolean set toTrueif the engine runs without network calls (line 89).supported_languages_hint: List of ISO-639-1 language codes for UI hints (line 92).
Abstract Methods to Implement
Three methods must be implemented to satisfy the contract:
is_available(cls) -> (bool, str)(line 95): Class method that performs sanity checks such as verifying API keys or dependencies exist, returning a tuple of availability status and message.generate(self, text, …) -> bytes(line 105): Instance method that synthesizes speech and returns raw audio bytes.list_voices(self) -> list[dict](line 127): Returns a list of voice metadata dictionaries containingid,name, andlanguagekeys used to populate the voice picker.
Optional Overrides
You may override get_sample_rate(self) -> int (line 135) if your engine outputs audio at a rate other than the default 24 kHz assumed by the VoiceStudio audio pipeline.
Plugin Registration and Discovery
Using the @register_plugin Decorator
Registration occurs through the @register_plugin decorator defined at line 35 in backend/services/plugin_sdk.py. Applying this decorator to your class validates the presence of the id attribute and adds the plugin to the global PLUGINS registry. Alternatively, you can manually insert your class into the PLUGINS dictionary, though the decorator is the recommended approach.
Automatic Discovery Mechanism
VoiceStudio automatically imports all Python files in the backend/plugins/ directory at startup via the discover_plugins() function (line 50). This means placing your file at backend/plugins/myengine.py is sufficient for the system to load and register your plugin without additional configuration or modifications to core application files.
Minimal Implementation Skeleton
# backend/plugins/myengine.py
from services.plugin_sdk import TTSPlugin, register_plugin
@register_plugin
class MyEnginePlugin(TTSPlugin):
"""MyEngine – a hypothetical local TTS engine."""
id = "myengine"
display_name = "MyEngine"
requires_api_key = False
is_local = True
supported_languages_hint = ["en", "es", "fr"]
@classmethod
def is_available(cls) -> tuple[bool, str]:
try:
import myengine # noqa: F401
return True, "Ready"
except ImportError:
return False, "pip install myengine"
def generate(self, text, *, voice_id=None, language=None, speed=1.0, **kw) -> bytes:
import myengine
audio = myengine.synthesize(text, voice=voice_id, lang=language, speed=speed)
return audio # bytes
def list_voices(self) -> list[dict]:
return [
{"id": "default", "name": "Default Voice", "language": "en"},
]
def get_sample_rate(self) -> int:
return 48000
Runtime Integration
Backend Instantiation
The backend retrieves plugin instances using the get_plugin(plugin_id) function from backend/services/plugin_sdk.py. This instantiates your concrete class and exposes the generate() method for API endpoints to call when processing synthesis requests.
Frontend Enumeration
The list_plugins() function provides JSON-compatible metadata to the frontend Settings UI. It consumes your class attributes and the is_available() return values to populate the TTS Engine picker, ensuring the interface remains responsive without importing heavy synthesis dependencies until actually needed.
Summary
- Inherit from
TTSPlugin: Subclass the abstract base class inbackend/services/plugin_sdk.pyand implement all required methods. - Define metadata: Set
id,display_name,requires_api_key,is_local, andsupported_languages_hintas class attributes. - Implement synthesis logic: Provide
is_available(),generate(), andlist_voices()methods with the exact signatures specified in the contract. - Register automatically: Place your file in
backend/plugins/and use the@register_plugindecorator to enable discovery. - Override defaults: Use
get_sample_rate()only if your engine outputs audio at a rate other than 24 kHz.
Frequently Asked Questions
What files do I need to modify to add a new TTS engine to VoiceStudio?
You only need to create a new Python file in the backend/plugins/ directory. The core application in backend/services/plugin_sdk.py handles the rest through automatic discovery, so no changes to existing code are required to register a new engine.
How does VoiceStudio detect my new TTS plugin automatically?
The discover_plugins() function in backend/services/plugin_sdk.py (line 50) automatically imports all .py files in backend/plugins/ when the module loads. If your class uses the @register_plugin decorator, it gets added to the global registry immediately upon import.
What is the difference between local and cloud TTS engines in the SDK?
The is_local boolean attribute distinguishes engines that run entirely on the host machine (set to True) from those requiring network calls to external APIs (set to False). The requires_api_key attribute works similarly to indicate whether the engine needs authentication credentials, which triggers UI prompts for API key configuration.
Can I override the default audio sample rate in a VoiceStudio TTS plugin?
Yes. The default sample rate is 24 kHz, but you can override the get_sample_rate() method (line 135 in backend/services/plugin_sdk.py) to return any integer value such as 48000 or 22050, ensuring the audio pipeline processes your engine's output correctly.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →