How to Configure Different STT Providers in LiveKit Agents
Configure different STT providers in LiveKit Agents by installing the specific plugin, setting the vendor API key via environment variable or constructor argument, and instantiating the provider's STT class or using the livekit.agents.inference.STT wrapper with a model string like "deepgram/nova-3:en".
LiveKit Agents provides a unified framework for building voice AI applications, abstracting speech-to-text functionality behind a common interface. To configure different STT providers such as Deepgram, Google Cloud Speech, or AssemblyAI, you instantiate provider-specific classes from their respective plugins or use the LiveKit Cloud Inference wrapper for dynamic provider selection. This guide covers the architecture, configuration patterns, and concrete implementation details derived from the livekit/agents source code.
Architecture Overview
LiveKit Agents abstracts STT behind the livekit.agents.stt.STT interface defined in livekit-agents/livekit/agents/stt/stt.py. This abstract base class declares the public API methods recognize and stream, along with capabilities such as streaming, interim_results, and diarization.
Provider plugins live in the livekit-plugins directory and contain concrete implementations that subclass the base STT class. For example, livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py implements _recognize_impl and stream for Deepgram's API.
The livekit.agents.inference.STT wrapper, located in livekit-agents/livekit/agents/inference/stt.py, offers a provider-agnostic entry point. It parses model strings (e.g., "deepgram/nova-3:en") and instantiates the appropriate provider class behind the scenes, also supporting fallback model chains.
Configuration Steps
Regardless of which provider you choose, the configuration pattern follows these steps:
- Install the plugin package using pip (e.g.,
pip install livekit-agents[deepgram]or install the full package). - Provide the vendor API key either by passing
api_keyexplicitly to the constructor or by setting the provider-specific environment variable (e.g.,DEEPGRAM_API_KEY,GOOGLE_APPLICATION_CREDENTIALS,ASSEMBLYAI_API_KEY). - Select a model using the provider's model identifier literals (e.g.,
"nova-3","latest_long"). - Configure options via the provider's
STTOptionsdataclass, passed throughextra_kwargsor direct parameters (e.g.,interim_results,punctuate,sample_rate). - (Optional) Define fallback models as a list of alternative providers to enable automatic failover on errors.
Provider-Specific Setup
Deepgram
Deepgram integration resides in livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py. The class expects a DeepgramOptions dataclass that controls features like interim_results, punctuate, and endpointing_ms.
from livekit.plugins.deepgram import stt as deepgram_stt
dg = deepgram_stt.STT(
model="nova-3",
language="en-US",
api_key="YOUR_DEEPGRAM_API_KEY", # Or set DEEPGRAM_API_KEY env var
extra_kwargs={
"interim_results": True,
"punctuate": True,
"endpointing_ms": 25,
},
fallback=[{"model": "assemblyai/universal-streaming"}],
)
stream = dg.stream()
# Feed audio buffers and await stream.next()
Google Cloud Speech
The Google provider in livekit-plugins/livekit-plugins-google/livekit/plugins/google/stt.py supports multiple languages and service account credentials via GOOGLE_APPLICATION_CREDENTIALS.
from livekit.plugins.google import stt as google_stt
gs = google_stt.STT(
languages=["en-US", "es-ES"],
detect_language=False,
interim_results=True,
punctuate=True,
enable_word_time_offsets=True,
model="latest_long",
# Credentials loaded from GOOGLE_APPLICATION_CREDENTIALS automatically
)
stream = gs.stream()
# Yields SpeechEvent objects with word-level timestamps
AssemblyAI
AssemblyAI's implementation in livekit-plugins/livekit-plugins-assemblyai/livekit/plugins/assemblyai/stt.py is streaming-only and exposes VAD threshold controls.
from livekit.plugins.assemblyai import stt as aai_stt
aai = aai_stt.STT(
api_key="YOUR_ASSEMBLYAI_KEY", # Or set ASSEMBLYAI_API_KEY env var
model="universal-streaming-english",
language_detection=False,
vad_threshold=0.4,
min_turn_silence=150, # milliseconds of silence before end-of-turn
)
stream = aai.stream()
# Events contain INTERIM_TRANSCRIPT or FINAL_TRANSCRIPT types
Cartesia
Cartesia provides a Whisper-based STT in livekit-plugins/livekit-plugins-cartesia/livekit/plugins/cartesia/stt.py.
from livekit.plugins.cartesia import stt as cartesia_stt
cs = cartesia_stt.STT(
model="ink-whisper",
api_key="YOUR_CARTESIA_KEY",
extra_kwargs={"min_volume": -0.5},
)
stream = cs.stream()
ElevenLabs
ElevenLabs offers real-time transcription via Scribe v2, implemented in livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/stt.py.
from livekit.plugins.elevenlabs import stt as eleven_stt
el = eleven_stt.STT(
model="scribe_v2_realtime",
api_key="YOUR_ELEVENLABS_KEY",
extra_kwargs={"include_timestamps": True},
)
stream = el.stream()
LiveKit Cloud Inference Wrapper
For a provider-agnostic approach, use livekit.agents.inference.STT from livekit-agents/livekit/agents/inference/stt.py. This wrapper parses model strings like "deepgram/nova-3:en" and supports automatic fallback chains.
from livekit.agents.inference import STT
stt_client = STT(
model="deepgram/nova-3:en",
fallback=[
{"model": "assemblyai/universal-streaming"},
{"model": "google/latest_long"},
],
# API keys are read from standard environment variables (DEEPGRAM_API_KEY, etc.)
)
stream = stt_client.stream()
The wrapper resolves the model string by mapping the provider segment to the corresponding plugin class, extracting the model name and optional language code, then instantiating the concrete provider with the appropriate STTOptions.
Complete Working Example
The following example demonstrates configuring a multi-provider STT pipeline with fallback support:
import os
import asyncio
from livekit.agents.inference import STT
# Set environment variables (or pass api_key directly to providers)
os.environ["DEEPGRAM_API_KEY"] = "dg_..."
os.environ["ASSEMBLYAI_API_KEY"] = "aa_..."
async def main():
# Configure primary Deepgram with AssemblyAI fallback
stt = STT(
model="deepgram/nova-3:en",
fallback=[
{"model": "assemblyai/universal-streaming"},
],
)
# Start streaming
stream = stt.stream()
# Process events (works identically regardless of which provider is active)
async for event in stream:
if event.type == "final_transcript":
print(f"Transcript: {event.alternatives[0].text}")
if __name__ == "__main__":
asyncio.run(main())
Summary
- Abstract Interface: All providers implement
livekit.agents.stt.STT, exposingrecognize()andstream()methods with consistentSpeechEventoutput. - Plugin Installation: Install individual provider packages (e.g.,
livekit-plugins-deepgram) or the full bundle. - Authentication: Provide API keys via environment variables (
DEEPGRAM_API_KEY,GOOGLE_APPLICATION_CREDENTIALS,ASSEMBLYAI_API_KEY) or explicit constructor arguments. - Model Selection: Use provider-specific model literals (e.g.,
"nova-3","latest_long") or the inference wrapper's model string format"provider/model:language". - Fallback Support: Define alternative providers in the
fallbackparameter to enable automatic failover during outages. - Source Locations: Core logic resides in
livekit-agents/livekit/agents/stt/stt.py(base class) andlivekit-agents/livekit/agents/inference/stt.py(wrapper), with provider implementations inlivekit-plugins/livekit-plugins-<provider>/livekit/plugins/<provider>/stt.py.
Frequently Asked Questions
How do I switch between STT providers without changing my application code?
Use the livekit.agents.inference.STT wrapper class. Pass a model string like "deepgram/nova-3:en" or "google/latest_long". The wrapper instantiates the correct provider plugin based on the string prefix, allowing you to change providers by modifying only the configuration string while keeping your streaming logic identical.
What environment variables do I need to set for authentication?
Each provider checks for specific environment variables: DEEPGRAM_API_KEY for Deepgram, GOOGLE_APPLICATION_CREDENTIALS (pointing to a service account JSON file) for Google Cloud Speech, and ASSEMBLYAI_API_KEY for AssemblyAI. You can override these by passing api_key or credentials_info directly to the constructor.
How does the fallback mechanism work when a provider fails?
When you provide a fallback list to STT or the inference wrapper, LiveKit attempts to instantiate the primary provider first. If that provider raises an error during initialization or streaming, the framework automatically iterates through the fallback list, instantiating each alternative provider in order until one succeeds. Each fallback entry accepts a model string and optional extra_kwargs for provider-specific configuration.
Can I use multiple languages or auto-detection with these providers?
Yes. Google Cloud Speech supports multiple languages via the languages parameter (accepting a list of BCP-47 codes) and enables automatic language detection with detect_language=True. Deepgram supports auto-detection by setting language=None or specifying a language code like "en-US". AssemblyAI offers language_detection as a boolean option in its configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →