How to Implement Video Avatar Integration with LiveKit Agents: Tavus, Hedra, and More
LiveKit Agents provides a modular avatar plugin system that lets you attach realtime video avatars like Tavus, Hedra, TruGen, LiveAvatar, and Simli to any voice agent by subclassing AvatarSession and bridging audio through DataStreamAudioOutput or QueueAudioOutput.
The livekit/agents repository offers a production-ready framework for building multimodal AI agents. Video avatar integration allows these agents to stream synchronized audio and video through a virtual participant, creating lifelike conversational experiences without managing complex media pipelines yourself.
Architecture of the Avatar Plugin System
The avatar system is built around a base class that standardizes how providers connect to a LiveKit room, publish media, and handle lifecycle events.
Core Components
The foundation resides in livekit/agents/voice/avatar.py, which defines the AvatarSession base class. This class manages:
- Virtual Participant Creation: Automatically joins the room under a distinct identity (e.g.,
tavus-avatar-agent) to publish audio and video tracks separately from the user or main agent. - Audio Bridging: Accepts PCM audio chunks from the TTS pipeline via
DataStreamAudioOutputorQueueAudioOutputand forwards them to the avatar provider for lip-sync. - Lifecycle Management: Monitors connection health, implements idle timeouts (default 30 seconds), and handles provider reconnections.
Provider Implementations
Each avatar service ships as an independent plugin under livekit-plugins/. All follow the same structural pattern:
| Provider | Plugin Package | Core Session File | Identity Constant |
|---|---|---|---|
| Tavus | livekit-plugins-tavus |
livekit/plugins/tavus/avatar.py |
_AVATAR_AGENT_IDENTITY = "tavus-avatar-agent" |
| Hedra | livekit-plugins-hedra |
livekit/plugins/hedra/avatar.py |
_AVATAR_AGENT_IDENTITY = "hedra-avatar-agent" |
| TruGen | livekit-plugins-trugen |
livekit/plugins/trugen/avatar.py |
_AVATAR_AGENT_IDENTITY = "trugen-avatar" |
| LiveAvatar | livekit-plugins-liveavatar |
livekit/plugins/liveavatar/avatar.py |
_AVATAR_AGENT_IDENTITY = "liveavatar-avatar-agent" |
| Simli | livekit-plugins-simli |
livekit/plugins/simli/avatar.py |
_AVATAR_AGENT_IDENTITY = "simli-avatar-agent" |
Each implementation subclasses AvatarSession and overrides _connect() for provider-specific handshakes and _run() to handle the audio-to-video streaming loop.
Implementing a Video Avatar Integration
Integrating a provider like Tavus or Hedra requires three steps: installing the plugin, instantiating the session, and wiring it to your agent's audio pipeline.
Installation and Setup
Install the specific provider plugin alongside the core agents package:
# For Tavus
pip install livekit-plugins-tavus
# For Hedra
pip install livekit-plugins-hedra
# For TruGen, LiveAvatar, or Simli, use the respective package name
Set your provider API key as an environment variable or pass it directly to the session constructor.
Creating the Avatar Session
Instantiate the provider-specific AvatarSession within your agent initialization. This example uses Tavus:
from livekit.plugins.tavus import AvatarSession as TavusAvatar
from livekit.agents.voice import Agent, VoiceConfig
class MyAgent(Agent):
def __init__(self):
super().__init__(voice=VoiceConfig())
# Initialize the avatar with provider-specific settings
self.avatar = TavusAvatar(
avatar_id="my-tavus-avatar-id",
avatar_participant_name="Guide",
avatar_participant_identity="tavus-avatar-agent"
)
# Attach the avatar's audio input to the agent's output pipeline
self.add_audio_output(self.avatar.audio_output)
The avatar_id corresponds to your configured avatar within the provider's dashboard. The avatar_participant_identity must be unique in the room and typically follows the pattern {provider}-avatar-agent as defined in the source files.
Wiring Audio to the Agent
The AgentSession class in livekit/agents/voice/agent_session.py orchestrates the LLM-to-TTS flow. When you call ctx.speak(response), the synthesized audio automatically routes to all attached audio outputs, including your avatar session:
async def on_message(self, ctx):
# Process user input through your LLM
response = await self.llm.run(ctx.transcript)
# This streams through TTS -> avatar.audio_output -> Tavus/Hedra video feed
await ctx.speak(response)
The avatar session receives PCM chunks via DataStreamAudioOutput or QueueAudioOutput, timestamps them, and forwards them to the provider's streaming endpoint for real-time lip-sync generation.
Advanced Configuration and Lifecycle Management
The avatar system provides hooks for complex interaction patterns and resilience.
Dynamic Avatar Switching
You can instantiate multiple AvatarSession objects and swap them based on conversation context:
# Initialize multiple avatars
self.sales_avatar = TavusAvatar(avatar_id="sales-rep")
self.support_avatar = TavusAvatar(avatar_id="support-agent")
# Switch based on intent
if intent == "sales":
await self.support_avatar.stop()
await self.sales_avatar.start(ctx.room)
self.add_audio_output(self.sales_avatar.audio_output)
Each avatar maintains its own participant identity, allowing seamless handoffs without disrupting the main agent session.
Idle Timeout and Reconnection
All built-in sessions expose max_idle_seconds (default 30 seconds) to automatically disconnect inactive avatars:
self.avatar = TavusAvatar(
avatar_id="my-avatar",
max_idle_seconds=60 # Keep alive for 1 minute of silence
)
The _run() method in livekit/agents/voice/avatar.py monitors the audio queue health. If the provider's WebSocket drops, the session attempts reconnection using exponential backoff before finally calling _disconnect() to clean up the virtual participant.
Summary
- LiveKit Agents provides a modular avatar plugin system in
livekit/agents/voice/avatar.pythat standardizes how video avatars connect to voice agents. - Each provider (Tavus, Hedra, TruGen, LiveAvatar, Simli) ships as a separate package under
livekit-plugins/with its ownAvatarSessionimplementation and unique participant identity constant. - Integration requires three steps: install the plugin, instantiate the provider-specific session with an
avatar_id, and attachsession.audio_outputto your agent usingadd_audio_output(). - The audio bridge automatically streams synthesized TTS audio to the avatar for real-time lip-sync via
DataStreamAudioOutputorQueueAudioOutput. - Lifecycle management includes configurable idle timeouts (
max_idle_seconds), automatic reconnection logic, and clean virtual participant teardown.
Frequently Asked Questions
How do I choose between Tavus, Hedra, and other avatar providers in LiveKit Agents?
Select based on your latency requirements and avatar customization needs. Tavus and LiveAvatar excel at sub-second latency for real-time conversations, while Hedra and TruGen offer more stylized or 3D avatar options. All providers follow the same AvatarSession interface in livekit/agents/voice/avatar.py, so you can swap implementations by changing the import statement and avatar_id without restructuring your agent code.
Can I run multiple video avatars simultaneously in the same LiveKit room?
Yes, by instantiating separate AvatarSession objects with unique participant identities. Each avatar publishes as a distinct participant (e.g., tavus-avatar-agent and simli-avatar-agent) defined by the _AVATAR_AGENT_IDENTITY constant in each plugin. Attach both audio_output instances to your agent, or route specific utterances to specific avatars using conditional logic in your on_message handler.
What happens if the avatar provider's connection drops during a conversation?
The AvatarSession automatically attempts reconnection with exponential backoff before gracefully disconnecting. The _run() method in livekit/agents/voice/avatar.py monitors the provider's WebSocket or HTTP stream health. If the connection fails, the session enters a reconnection loop. If reconnection fails permanently or the max_idle_seconds timeout elapses without audio activity, the session calls _disconnect() to unpublish tracks and remove the virtual participant from the room.
How do I synchronize the avatar's lip movements with the TTS audio output?
Lip-sync happens automatically when you attach the avatar's audio_output to your agent's voice pipeline. The AvatarSession receives timestamped PCM audio chunks via DataStreamAudioOutput or QueueAudioOutput and forwards them to the provider's streaming endpoint. Providers like Tavus, TruGen, and LiveAvatar use these audio timestamps to generate corresponding video frames with synchronized lip movements. No additional synchronization code is required in your agent implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →