How to Configure Different STT Providers in LiveKit Agents

Configure different STT providers in LiveKit Agents by installing the specific plugin, setting the vendor API key via environment variable or constructor argument, and instantiating the provider's STT class or using the livekit.agents.inference.STT wrapper with a model string like "deepgram/nova-3:en".

LiveKit Agents provides a unified framework for building voice AI applications, abstracting speech-to-text functionality behind a common interface. To configure different STT providers such as Deepgram, Google Cloud Speech, or AssemblyAI, you instantiate provider-specific classes from their respective plugins or use the LiveKit Cloud Inference wrapper for dynamic provider selection. This guide covers the architecture, configuration patterns, and concrete implementation details derived from the livekit/agents source code.

Architecture Overview

LiveKit Agents abstracts STT behind the livekit.agents.stt.STT interface defined in livekit-agents/livekit/agents/stt/stt.py. This abstract base class declares the public API methods recognize and stream, along with capabilities such as streaming, interim_results, and diarization.

Provider plugins live in the livekit-plugins directory and contain concrete implementations that subclass the base STT class. For example, livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py implements _recognize_impl and stream for Deepgram's API.

The livekit.agents.inference.STT wrapper, located in livekit-agents/livekit/agents/inference/stt.py, offers a provider-agnostic entry point. It parses model strings (e.g., "deepgram/nova-3:en") and instantiates the appropriate provider class behind the scenes, also supporting fallback model chains.

Configuration Steps

Regardless of which provider you choose, the configuration pattern follows these steps:

  1. Install the plugin package using pip (e.g., pip install livekit-agents[deepgram] or install the full package).
  2. Provide the vendor API key either by passing api_key explicitly to the constructor or by setting the provider-specific environment variable (e.g., DEEPGRAM_API_KEY, GOOGLE_APPLICATION_CREDENTIALS, ASSEMBLYAI_API_KEY).
  3. Select a model using the provider's model identifier literals (e.g., "nova-3", "latest_long").
  4. Configure options via the provider's STTOptions dataclass, passed through extra_kwargs or direct parameters (e.g., interim_results, punctuate, sample_rate).
  5. (Optional) Define fallback models as a list of alternative providers to enable automatic failover on errors.

Provider-Specific Setup

Deepgram

Deepgram integration resides in livekit-plugins/livekit-plugins-deepgram/livekit/plugins/deepgram/stt.py. The class expects a DeepgramOptions dataclass that controls features like interim_results, punctuate, and endpointing_ms.

from livekit.plugins.deepgram import stt as deepgram_stt

dg = deepgram_stt.STT(
    model="nova-3",
    language="en-US",
    api_key="YOUR_DEEPGRAM_API_KEY",  # Or set DEEPGRAM_API_KEY env var

    extra_kwargs={
        "interim_results": True,
        "punctuate": True,
        "endpointing_ms": 25,
    },
    fallback=[{"model": "assemblyai/universal-streaming"}],
)

stream = dg.stream()

# Feed audio buffers and await stream.next()

Google Cloud Speech

The Google provider in livekit-plugins/livekit-plugins-google/livekit/plugins/google/stt.py supports multiple languages and service account credentials via GOOGLE_APPLICATION_CREDENTIALS.

from livekit.plugins.google import stt as google_stt

gs = google_stt.STT(
    languages=["en-US", "es-ES"],
    detect_language=False,
    interim_results=True,
    punctuate=True,
    enable_word_time_offsets=True,
    model="latest_long",
    # Credentials loaded from GOOGLE_APPLICATION_CREDENTIALS automatically

)

stream = gs.stream()

# Yields SpeechEvent objects with word-level timestamps

AssemblyAI

AssemblyAI's implementation in livekit-plugins/livekit-plugins-assemblyai/livekit/plugins/assemblyai/stt.py is streaming-only and exposes VAD threshold controls.

from livekit.plugins.assemblyai import stt as aai_stt

aai = aai_stt.STT(
    api_key="YOUR_ASSEMBLYAI_KEY",  # Or set ASSEMBLYAI_API_KEY env var

    model="universal-streaming-english",
    language_detection=False,
    vad_threshold=0.4,
    min_turn_silence=150,  # milliseconds of silence before end-of-turn

)

stream = aai.stream()

# Events contain INTERIM_TRANSCRIPT or FINAL_TRANSCRIPT types

Cartesia

Cartesia provides a Whisper-based STT in livekit-plugins/livekit-plugins-cartesia/livekit/plugins/cartesia/stt.py.

from livekit.plugins.cartesia import stt as cartesia_stt

cs = cartesia_stt.STT(
    model="ink-whisper",
    api_key="YOUR_CARTESIA_KEY",
    extra_kwargs={"min_volume": -0.5},
)

stream = cs.stream()

ElevenLabs

ElevenLabs offers real-time transcription via Scribe v2, implemented in livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/stt.py.

from livekit.plugins.elevenlabs import stt as eleven_stt

el = eleven_stt.STT(
    model="scribe_v2_realtime",
    api_key="YOUR_ELEVENLABS_KEY",
    extra_kwargs={"include_timestamps": True},
)

stream = el.stream()

LiveKit Cloud Inference Wrapper

For a provider-agnostic approach, use livekit.agents.inference.STT from livekit-agents/livekit/agents/inference/stt.py. This wrapper parses model strings like "deepgram/nova-3:en" and supports automatic fallback chains.

from livekit.agents.inference import STT

stt_client = STT(
    model="deepgram/nova-3:en",
    fallback=[
        {"model": "assemblyai/universal-streaming"},
        {"model": "google/latest_long"},
    ],
    # API keys are read from standard environment variables (DEEPGRAM_API_KEY, etc.)

)

stream = stt_client.stream()

The wrapper resolves the model string by mapping the provider segment to the corresponding plugin class, extracting the model name and optional language code, then instantiating the concrete provider with the appropriate STTOptions.

Complete Working Example

The following example demonstrates configuring a multi-provider STT pipeline with fallback support:

import os
import asyncio
from livekit.agents.inference import STT

# Set environment variables (or pass api_key directly to providers)

os.environ["DEEPGRAM_API_KEY"] = "dg_..."
os.environ["ASSEMBLYAI_API_KEY"] = "aa_..."

async def main():
    # Configure primary Deepgram with AssemblyAI fallback

    stt = STT(
        model="deepgram/nova-3:en",
        fallback=[
            {"model": "assemblyai/universal-streaming"},
        ],
    )
    
    # Start streaming

    stream = stt.stream()
    
    # Process events (works identically regardless of which provider is active)

    async for event in stream:
        if event.type == "final_transcript":
            print(f"Transcript: {event.alternatives[0].text}")

if __name__ == "__main__":
    asyncio.run(main())

Summary

  • Abstract Interface: All providers implement livekit.agents.stt.STT, exposing recognize() and stream() methods with consistent SpeechEvent output.
  • Plugin Installation: Install individual provider packages (e.g., livekit-plugins-deepgram) or the full bundle.
  • Authentication: Provide API keys via environment variables (DEEPGRAM_API_KEY, GOOGLE_APPLICATION_CREDENTIALS, ASSEMBLYAI_API_KEY) or explicit constructor arguments.
  • Model Selection: Use provider-specific model literals (e.g., "nova-3", "latest_long") or the inference wrapper's model string format "provider/model:language".
  • Fallback Support: Define alternative providers in the fallback parameter to enable automatic failover during outages.
  • Source Locations: Core logic resides in livekit-agents/livekit/agents/stt/stt.py (base class) and livekit-agents/livekit/agents/inference/stt.py (wrapper), with provider implementations in livekit-plugins/livekit-plugins-<provider>/livekit/plugins/<provider>/stt.py.

Frequently Asked Questions

How do I switch between STT providers without changing my application code?

Use the livekit.agents.inference.STT wrapper class. Pass a model string like "deepgram/nova-3:en" or "google/latest_long". The wrapper instantiates the correct provider plugin based on the string prefix, allowing you to change providers by modifying only the configuration string while keeping your streaming logic identical.

What environment variables do I need to set for authentication?

Each provider checks for specific environment variables: DEEPGRAM_API_KEY for Deepgram, GOOGLE_APPLICATION_CREDENTIALS (pointing to a service account JSON file) for Google Cloud Speech, and ASSEMBLYAI_API_KEY for AssemblyAI. You can override these by passing api_key or credentials_info directly to the constructor.

How does the fallback mechanism work when a provider fails?

When you provide a fallback list to STT or the inference wrapper, LiveKit attempts to instantiate the primary provider first. If that provider raises an error during initialization or streaming, the framework automatically iterates through the fallback list, instantiating each alternative provider in order until one succeeds. Each fallback entry accepts a model string and optional extra_kwargs for provider-specific configuration.

Can I use multiple languages or auto-detection with these providers?

Yes. Google Cloud Speech supports multiple languages via the languages parameter (accepting a list of BCP-47 codes) and enables automatic language detection with detect_language=True. Deepgram supports auto-detection by setting language=None or specifying a language code like "en-US". AssemblyAI offers language_detection as a boolean option in its configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →