How to Configure Different TTS Providers in LiveKit Agents: A Complete Guide
LiveKit Agents provides a unified TTS class that configures any supported provider using a model string format <provider>/<model>[:<voice]> and automatically handles authentication through environment variables or constructor arguments.
The livekit/agents repository offers a provider-agnostic abstraction for text-to-speech synthesis. Whether you need OpenAI's GPT-4o-mini-TTS, ElevenLabs' multilingual models, or Cartesia's streaming voices, you configure them through a single consistent interface defined in livekit-agents/livekit/agents/inference/tts.py.
Understanding the Unified TTS Abstraction
Core TTS Class and Model String Parsing
The TTS class constructor in livekit/agents/inference/tts.py accepts a model string that determines which provider to use. The parser expects the format:
<provider>/<model>[:<voice>]
The _parse_model_string method (lines 64-76) splits this identifier to route requests to the correct plugin. For example, cartesia/sonic-2:en_us_001 instructs the system to use the Cartesia plugin with the Sonic-2 model and a specific US English voice.
Environment Variables and Authentication
The unified layer automatically reads standard environment variables for authentication. As implemented in lines 30-48 of tts.py, the system checks for:
LIVEKIT_INFERENCE_URLLIVEKIT_INFERENCE_API_KEYLIVEKIT_INFERENCE_API_SECRET
Provider-specific API keys follow standard naming conventions (e.g., OPENAI_API_KEY, ELEVENLABS_API_KEY, CARTESIA_API_KEY). You can pass these explicitly via the constructor or rely on environment detection.
Supported TTS Providers and Model Strings
Each provider implements a subclass under livekit-plugins/. The following table maps providers to their plugin paths and typical model strings:
| Provider | Plugin Path | Example Model String | Default Model |
|---|---|---|---|
| OpenAI | livekit-plugins-openai/livekit/plugins/openai/tts.py |
openai/gpt-4o-mini-tts:ash |
gpt-4o-mini-tts |
| ElevenLabs | livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/tts.py |
elevenlabs/eleven_multilingual_v2:en_us_001 |
eleven_multilingual_v2 |
| Cartesia | livekit-plugins-cartesia/livekit/plugins/cartesia/tts.py |
cartesia/sonic:en_us_001 |
sonic |
| Deepgram | livekit-plugins-deepgram/livekit/plugins/deepgram/tts.py |
deepgram/aura:en-US |
aura |
| Rime | livekit-plugins-rime/livekit/plugins/rime/tts.py |
rime/arcana:en |
arcana |
| Inworld | livekit-plugins-inworld/livekit/plugins/inworld/tts.py |
inworld/inworld-tts-1:en-US |
inworld-tts-1 |
Each plugin defines its own options dataclass (e.g., OpenAI._TTSOptions) and implements the synthesize method to communicate with provider HTTP APIs.
Configuring TTS Providers in Practice
Basic Configuration with Cartesia
To configure a specific provider, instantiate the unified TTS class with the appropriate model string and optional extra_kwargs for provider-specific settings.
from livekit.agents import inference
# Configure Cartesia with specific voice and emotional tone
tts = inference.TTS(
model="cartesia/sonic-2",
voice="en_us_001",
extra_kwargs={"emotion": "happy", "speed": "fast"},
)
# Pre-warm connection pool to reduce first-call latency
tts.prewarm()
# Synthesize audio
audio_stream = tts.synthesize("Hello, LiveKit agents!")
async for chunk in audio_stream:
# chunk contains raw PCM data (16-bit, 24kHz)
process_audio(chunk)
The constructor handles model string parsing and option validation (lines 61-78 in tts.py).
Implementing Fallback Chains
For production reliability, configure automatic fallback to alternative providers using the fallback parameter. The _normalize_fallback method (lines 100-112) processes fallback specifications.
from livekit.agents import inference
tts = inference.TTS(
model="openai/gpt-4o-mini-tts",
voice="ash",
fallback=[
"elevenlabs/eleven_multilingual_v2:en_us_001",
{
"model": "deepgram/aura",
"voice": "en-US",
"extra_kwargs": {"speed": 1.2}
},
],
conn_options=inference.APIConnectOptions(timeout=10, max_retry=2),
)
# If OpenAI fails, automatically tries ElevenLabs, then Deepgram
audio = await tts.synthesize("Fallback demonstration").read()
Direct Provider Instantiation
For advanced scenarios requiring provider-specific configuration (such as Azure OpenAI endpoints), instantiate the plugin class directly rather than using the unified interface.
from livekit_plugins.openai import tts as openai_tts
# Direct Azure OpenAI configuration
client = openai_tts.TTS.with_azure(
model="gpt-4o-mini-tts",
voice="ash",
azure_endpoint="https://my-azure.openai.azure.com",
api_key="my-azure-key",
)
audio = await client.synthesize("Direct provider instantiation").read()
The with_azure factory method is defined in livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/tts.py (lines 150-190).
Connection and Performance Options
Control retry behavior and timeouts using APIConnectOptions. Pass a custom instance via conn_options or accept the default DEFAULT_API_CONNECT_OPTIONS.
from livekit.agents import inference
conn_opts = inference.APIConnectOptions(
timeout=30.0,
max_retry=3
)
tts = inference.TTS(
model="elevenlabs/eleven_multilingual_v2",
conn_options=conn_opts
)
All providers expose three high-level methods:
synthesize(text, conn_options=...)– Returns aChunkedStreamfor async iteration over audio bytes.stream(conn_options=...)– Returns aSynthesizeStreamfor real-time token-by-token generation (supported by streaming providers like Cartesia).prewarm()– Opens WebSocket pools in advance to reduce first-call latency.
Summary
- LiveKit Agents provides a unified
TTSclass inlivekit/agents/inference/tts.pythat configures any supported provider using a model string format<provider>/<model>[:<voice>]. - Authentication works via environment variables (
OPENAI_API_KEY,ELEVENLABS_API_KEY, etc.) or direct constructor arguments. - Provider-specific options pass through the
extra_kwargsparameter, while fallback chains ensure reliability via thefallbackparameter processed by_normalize_fallback. - Direct instantiation of plugin classes (e.g.,
livekit_plugins.openai.tts.TTS) enables advanced configurations like Azure OpenAI endpoints. - Performance tuning uses
APIConnectOptionsto control timeouts and retries, withprewarm()available to reduce cold-start latency.
Frequently Asked Questions
How do I switch between TTS providers without changing my code?
Use the unified inference.TTS class with different model strings. The constructor parses strings like openai/gpt-4o-mini-tts or elevenlabs/eleven_multilingual_v2 and automatically loads the correct plugin. Your synthesis code remains identical regardless of the provider.
What environment variables do I need for each TTS provider?
Each provider expects its standard API key environment variable: OPENAI_API_KEY for OpenAI, ELEVENLABS_API_KEY for ElevenLabs, CARTESIA_API_KEY for Cartesia, and DEEPGRAM_API_KEY for Deepgram. The unified TTS layer also respects LIVEKIT_INFERENCE_URL, LIVEKIT_INFERENCE_API_KEY, and LIVEKIT_INFERENCE_API_SECRET for custom inference endpoints.
How do I configure Azure OpenAI instead of the standard OpenAI endpoint?
Instantiate the OpenAI plugin directly using the with_azure factory method rather than the unified TTS class. Pass your azure_endpoint, api_key, and deployment model to livekit_plugins.openai.tts.TTS.with_azure(). This bypasses the standard model-string parsing and connects directly to your Azure OpenAI resource.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →