How to Optimize Latency and Performance in LiveKit Agents: A Complete Guide
Enable preemptive generation, use balanced latency mode in TTS plugins, and tighten endpointing delays to reduce end-to-end latency from over 1 second to approximately 500ms.
The LiveKit Agents framework powers real-time voice AI applications by orchestrating speech-to-text, LLM inference, and text-to-synthesis pipelines. To deliver conversational experiences that feel instantaneous, developers must optimize latency and performance at every stage of this pipeline. This guide examines the specific configuration options, source code locations, and implementation patterns that minimize round-trip delay in the LiveKit Agents repository.
Understanding the LiveKit Agents Pipeline Architecture
LiveKit Agents processes audio through a sequential pipeline: audio capture → Voice Activity Detection (VAD) → Speech-to-Text (STT) → turn detection → Large Language Model (LLM) inference → Text-to-Speech (TTS) → network transport. Latency accumulates at each handoff point. The framework provides specific knobs in AgentSession and individual plugins to parallelize work, reduce waiting periods, and reuse connections.
Key Strategies to Optimize Latency and Performance
Enable Preemptive Generation
Preemptive generation overlaps LLM inference with user speech. When enabled, the model begins generating a response as soon as partial transcripts arrive, rather than waiting for the user to stop speaking. According to test suites in the repository, this reduces end-to-end latency from approximately 1.1 seconds to 0.8 seconds.
In livekit-agents/livekit/agents/voice/agent_session.py, set the preemptive_generation parameter:
session = AgentSession(
preemptive_generation=True,
)
The underlying implementation creates a _preemptive_generation object in AgentActivity at line 1437 of livekit-agents/livekit/agents/voice/agent_activity.py, which is cancelled if the user speaks again (lines 975-978).
Configure TTS Latency Modes
Streaming TTS plugins expose latency modes that trade quality for speed. Use balanced mode for the lowest round-trip time.
For FishAudio, configure latency_mode in livekit-plugins/livekit-plugins-fishaudio/livekit/plugins/fishaudio/tts.py:
from livekit.plugins.fishaudio import FishAudioTTS
tts = FishAudioTTS(
api_key="YOUR_API_KEY",
latency_mode="balanced", # ~300ms round-trip
)
For ElevenLabs, enable auto_mode in livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/tts.py to disable chunk scheduling and synthesize sentence-by-sentence:
from livekit.plugins.elevenlabs import ElevenLabsTTS
tts = ElevenLabsTTS(
voice_id="YOUR_VOICE_ID",
auto_mode=True, # Reduces latency via sentence-level tokenization
streaming_latency=4, # Max latency optimization
)
Optimize Turn Detection and Endpointing
Tightening endpointing parameters reduces the silence duration required before the system processes a turn.
In livekit-agents/livekit/agents/voice/agent_session.py, configure AgentSessionOptions:
session = AgentSession(
turn_detection="vad", # Fast VAD-based detection
min_endpointing_delay=0.2, # Seconds of silence before ending turn
max_endpointing_delay=2.0, # Maximum wait time
min_interruption_duration=0.1, # Ignore very short pauses
)
If your LLM plugin supports it, use turn_detection="realtime_llm" to eliminate the separate turn-detection round-trip.
Reuse Connections with APIConnectOptions
All plugins accept APIConnectOptions to enable connection pooling and keep-alive behavior, reducing handshake overhead.
Defined in livekit/agents/types.py and used across plugins like elevenlabs/tts.py:
from livekit.agents import APIConnectOptions
conn_opts = APIConnectOptions(
timeout=5.0,
max_retries=2,
keepalive=True,
)
# Reuse connection across synthesis calls
audio_stream = tts.stream(conn_options=conn_opts)
Complete Implementation Examples
Low-Latency AgentSession Configuration
Combine all optimizations into a single high-performance setup:
from livekit.agents import AgentSession, APIConnectOptions
from livekit.plugins.fishaudio import FishAudioTTS
from livekit.plugins.deepgram import DeepgramSTT
# 1. Low-latency TTS with balanced mode
tts = FishAudioTTS(
api_key="FISH_API_KEY",
latency_mode="balanced",
)
# 2. Fast STT with tight endpointing
stt = DeepgramSTT(
api_key="DEEPGRAM_KEY",
endpointing_delay=0.2,
)
# 3. Session with preemptive generation and tight endpointing
session = AgentSession(
turn_detection="vad",
stt=stt,
tts=tts,
preemptive_generation=True,
min_endpointing_delay=0.2,
max_endpointing_delay=2.0,
allow_interruptions=True,
false_interruption_timeout=1.0,
)
# 4. Connect to room
session.connect(url="wss://my.livekit.server", token="my-token")
ElevenLabs Auto-Mode Setup
For ElevenLabs-specific optimizations:
from livekit.plugins.elevenlabs import ElevenLabsTTS
from livekit.agents import AgentSession
tts = ElevenLabsTTS(
voice_id="EXAMPLE_VOICE",
auto_mode=True, # Sentence-level tokenization
streaming_latency=4, # Max latency optimization
inactivity_timeout=60, # Keep connection alive
)
session = AgentSession(
tts=tts,
preemptive_generation=True,
turn_detection="realtime_llm", # If LLM supports it
)
Monitoring Latency with OpenTelemetry
Capture e2e_latency metrics for profiling:
from livekit.agents.telemetry import tracer
@tracer.start_as_current_span("voice_pipeline")
def run_pipeline():
# Your agent logic here
pass
# After execution, inspect the span attributes in Jaeger or Tempo
# The 'e2e_latency' metric appears under span attributes
The telemetry implementation in livekit/agents/telemetry/traces.py automatically attaches e2e_latency values to traces when available.
Critical Source Files for Performance Tuning
livekit-agents/livekit/agents/voice/agent_session.py– ContainsAgentSessionOptionsincludingpreemptive_generation,turn_detection, and endpointing delays.livekit-agents/livekit/agents/voice/agent_activity.py– Implements the_preemptive_generationlogic (creation at line 1437, cancellation at lines 975-978).livekit-plugins/livekit-plugins-fishaudio/livekit/plugins/fishaudio/tts.py– Exposeslatency_modeparameter (balanced/normal).livekit-plugins/livekit-plugins-elevenlabs/livekit/plugins/elevenlabs/tts.py– Providesauto_modeandstreaming_latencyconfiguration.livekit-plugins/livekit-plugins-ultravox/livekit/plugins/ultravox/realtime/realtime_model.py– Demonstrates latency measurement patterns for realtime models.livekit/agents/telemetry/traces.py– Emitse2e_latencymetrics via OpenTelemetry.livekit/agents/types.py– DefinesAPIConnectOptionsfor connection reuse across plugins.
Summary
- Enable preemptive generation in
AgentSessionto overlap LLM inference with user speech, reducing latency by approximately 300ms. - Select balanced latency mode in TTS plugins like FishAudio to achieve ~300ms round-trip times.
- Activate auto_mode for ElevenLabs to enable sentence-level tokenization and eliminate chunk scheduling delays.
- Tighten endpointing parameters (
min_endpointing_delay,max_endpointing_delay) to reduce silence detection time. - Reuse connections via
APIConnectOptionsto minimize handshake overhead across STT, LLM, and TTS calls. - Monitor
e2e_latencytraces inlivekit/agents/telemetry/traces.pyto identify bottlenecks iteratively.
Frequently Asked Questions
What is the typical latency reduction when enabling preemptive generation?
Enabling preemptive_generation=True typically reduces end-to-end latency from approximately 1.1 seconds to 0.8 seconds. This 300ms improvement occurs because the LLM begins generating responses while the user is still speaking, overlapping computation with audio capture.
How does the balanced latency mode affect TTS quality?
The latency_mode="balanced" setting in plugins like FishAudio prioritizes speed over maximum quality, delivering audio in approximately 300ms compared to 500ms in normal mode. While this may slightly reduce audio fidelity in complex acoustic scenarios, the difference is typically imperceptible for conversational voice applications.
Can I use realtime_llm turn detection with any LLM provider?
No, the turn_detection="realtime_llm" option requires specific support from your LLM plugin. This mode eliminates the separate turn-detection round-trip by having the LLM itself determine when the user has finished speaking. Check your specific plugin implementation in the livekit-plugins directory to verify compatibility.
Where can I view e2e_latency metrics in production?
The e2e_latency metrics are emitted via OpenTelemetry traces defined in livekit/agents/telemetry/traces.py. Configure a trace collector such as Jaeger or Grafana Tempo in your deployment, then inspect the span attributes for e2e_latency values to profile your pipeline performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →