How to Add a New TTS Provider to Muse Alongside ElevenLabs: A Complete Implementation Guide
To add a new TTS provider to the Muse app, implement the TTSProvider interface with three suspend functions, register your class in the Koin DI module, and optionally wrap it with GroupedTTSProvider for fallback support.
The Muse repository (kkoshin/muse) uses a clean abstraction layer that makes adding a new TTS provider straightforward without modifying existing UI or business logic. Whether you want to integrate a local speech engine, a different cloud API, or a mock for testing, the architecture allows you to add a new TTS provider alongside the default ElevenLabs integration through simple interface implementation and dependency injection.
Understanding the TTSProvider Contract
The core contract for text-to-speech functionality is defined in TTSProvider.kt located at muse/src/commonMain/kotlin/io/github/kkoshin/muse/core/provider/. Any class adding a new TTS provider must implement this interface with three suspend functions:
generate(voiceId: String, text: String): Result<TTSResult>– Produces an audio stream as anokio.Sourcewrapped in a result.queryQuota(): Result<CharacterQuota>– Returns remaining character limits or a dummy value if the service doesn't enforce quotas.queryVoices(): Result<List<Voice>>– Lists available voices with metadata including accent, age, gender, and preview URLs.
The default ElevenLabs implementation, ElevenLabProcessor, lives in muse/src/commonMain/kotlin/io/github/kkoshin/muse/core/manager/ and serves as a reference for how to structure network calls and error handling.
Step‑by‑Step Implementation Guide
Step 1: Create Your Custom Provider Class
Create a new Kotlin file in the vendor package that implements TTSProvider. The file location should mirror existing providers, such as muse/src/commonMain/kotlin/io/github/kkoshin/muse/tts/vendor/MyCustomTTSProvider.kt.
package io.github.kkoshin.muse.tts.vendor
import io.github.kkoshin.muse.core.provider.*
import io.github.kkoshin.muse.tts.*
class MyCustomTTSProvider(
private val apiKey: String
) : TTSProvider {
override suspend fun generate(voiceId: String, text: String): Result<TTSResult> {
// Implement your API call here
// Convert response bytes to okio.Source
// return Result.success(TTSResult(source, SupportedAudioType.MP3, metadata))
TODO("Replace with real implementation")
}
override suspend fun queryQuota(): Result<CharacterQuota> =
Result.success(CharacterQuota.empty)
override suspend fun queryVoices(): Result<List<Voice>> =
Result.success(listOf(
Voice(
voiceId = "custom-voice-1",
name = "Custom Neural Voice",
description = "High-quality synthetic voice",
previewUrl = "",
accent = Voice.Accent.Other,
age = Voice.Age.Other,
useCase = null,
gender = Voice.Gender.Other,
descriptive = null
)
))
}
For a concrete reference on structuring mock implementations or handling edge cases, examine MockTTSProvider.kt in muse/src/androidDebug/kotlin/io/github/kkoshin/muse/tts/vendor/.
Step 2: Register Your Provider in Koin
The dependency injection graph is configured in appModule.kt at muse/src/androidMain/kotlin/io/github/kkoshin/muse/. You have two registration strategies when adding a new TTS provider:
Option A: Replace the Default Provider
Use this approach if you want your implementation to be the sole TTS engine used throughout the app:
val appModule = module {
includes(baseModule)
single<TTSProvider> {
MyCustomTTSProvider(apiKey = "your-api-key-here")
}
// Keep other ElevenLabs services unchanged
single<AudioIsolationProvider> { ElevenLabProcessor(get(), get()) }
single<SoundEffectProvider> { ElevenLabProcessor(get(), get()) }
single<STTProvider> { ElevenLabProcessor(get(), get()) }
}
Option B: Combine Providers with GroupedTTSProvider
Use GroupedTTSProvider (located in muse/src/androidDebug/kotlin/io/github/kkoshin/muse/tts/vendor/) to enable fallback behavior or user selection between multiple engines:
val appModule = module {
includes(baseModule)
single<TTSProvider> {
GroupedTTSProvider(
providers = listOf(
MyCustomTTSProvider(apiKey = "your-key-here"),
ElevenLabProcessor(get(), get()) // fallback to ElevenLabs
)
)
}
single<AudioIsolationProvider> { ElevenLabProcessor(get(), get()) }
single<SoundEffectProvider> { ElevenLabProcessor(get(), get()) }
single<STTProvider> { ElevenLabProcessor(get(), get()) }
}
GroupedTTSProvider iterates over the provider list in order, returning the first successful result from generate() and aggregating quotas across all registered engines.
Step 3: Verify Injection Points
Once registered in Koin, your new provider automatically injects into all existing TTS consumers without code changes. Key injection points include:
SpeechProcessorManager– Orchestrates speech generation workflowsExportViewModel– Handles audio export functionality
Both classes declare dependencies on the TTSProvider interface, so Koin resolves them to your registered implementation at runtime.
Step 4: Optional UI Configuration
If you want users to toggle between providers, expose the provider list through a ProviderRegistry or store the active provider ID in MutableState persisted to DataStore. The GroupedTTSProvider already supports multiple providers; you can filter its internal list based on user preferences stored in your configuration layer.
Testing Your Implementation
The debug build provides a reference testing pattern through MockTTSProvider and its corresponding DI module mockAppModule.kt at muse/src/androidDebug/kotlin/io/github/kkoshin/muse/.
To test your new TTS provider without consuming real API quota:
- Create a test implementation that returns static
ByteArraydata ingenerate() - Register it in a test-specific Koin module
- Inject into
SpeechProcessorManagerto verify integration
The mock module demonstrates how to override production bindings with test doubles using Koin's module isolation.
Summary
- Implement
TTSProviderwithgenerate(),queryQuota(), andqueryVoices()methods to define your engine's contract. - Register in
appModule.ktusing either direct binding orGroupedTTSProviderfor multi-engine fallback support. - Leverage existing injection –
SpeechProcessorManagerandExportViewModelautomatically receive your implementation through Koin's interface resolution. - Reference
MockTTSProviderfor testing patterns and structure when adding a new TTS provider to the Muse codebase.
Frequently Asked Questions
What methods must I implement when adding a new TTS provider?
You must implement three suspend functions defined in the TTSProvider interface: generate(voiceId, text) to return a Result<TTSResult> containing the audio stream, queryQuota() to return character limits or CharacterQuota.empty, and queryVoices() to return a list of available Voice objects with metadata.
Can I use multiple TTS providers simultaneously in Muse?
Yes. Wrap your providers in GroupedTTSProvider, which attempts each engine in sequence until one succeeds. This allows you to add a new TTS provider as a primary option while keeping ElevenLabs as a fallback, or to offer user-selectable voice engines within the same session.
Where should I place my custom TTS provider files?
Follow the existing package structure by placing your implementation in muse/src/commonMain/kotlin/io/github/kkoshin/muse/tts/vendor/ for shared code, or muse/src/androidMain/... for Android-specific implementations. Mirror the file organization of MockTTSProvider.kt for consistency.
How do I test my new TTS provider without making real API calls?
Use the debug build's MockTTSProvider as a template to create a test double that returns static audio data or empty sources. Register your mock in a test-specific Koin module (similar to mockAppModule.kt) to override production bindings during unit or integration tests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →