How to Implement Background Audio Handling in LiveKit Agents: A Complete Guide
Use the BackgroundAudioPlayer class from livekit.agents.voice.background_audio to manage looping ambient sounds and "thinking" cues that trigger automatically when your agent enters the thinking state.
The livekit/agents repository provides a robust background audio subsystem that lets you add immersive ambient soundscapes and reactive audio cues to your voice agents. This implementation handles audio mixing, state-driven playback, and automatic interruption when the agent needs to speak.
What Is Background Audio Handling in LiveKit Agents?
Background audio handling in LiveKit Agents refers to the continuous playback of ambient sounds (like office noise or nature sounds) alongside reactive "thinking" cues that play when the LLM is processing. The system manages a single continuous local audio track published to the LiveKit room, mixing multiple audio sources through a Rust-backed rtc.AudioMixer for low-latency performance.
The BackgroundAudioPlayer class coordinates with AgentSession to automatically interrupt background sounds when the agent begins speaking or when a new user turn arrives.
Architecture Overview
The background audio system consists of several coordinated components:
| Component | Role | Key Methods | Source File |
|---|---|---|---|
BackgroundAudioPlayer |
Manages continuous audio track, ambient loops, and thinking cues | start(room, agent_session), play(source, loop=False), aclose() |
livekit/agents/voice/background_audio.py |
AudioConfig |
Wraps sound sources with volume and probability settings | AudioConfig(source, volume=1.0, probability=1.0) |
livekit/agents/voice/background_audio.py |
BuiltinAudioClip |
Enum of built-in OGG clips (office ambience, keyboard typing) | BuiltinAudioClip.OFFICE_AMBIENCE.path() |
livekit/agents/voice/background_audio.py |
PlayHandle |
Controls individual playback instances | stop(), done() |
livekit/agents/voice/background_audio.py |
AgentActivity |
Interrupts background audio on user turns | _interrupt_background_speeches() |
livekit/agents/voice/agent_activity.py |
Audio Flow:
BackgroundAudioPlayer.start()creates anrtc.AudioSourceandrtc.AudioMixer, publishing aLocalAudioTrackto the room- Ambient sounds loop continuously through the mixer
- When
AgentStateChangedEventfires withnew_state == "thinking", the player triggers the configured thinking sound AgentActivity._interrupt_background_speeches()stops thinking cues when the agent begins speaking
Step-by-Step Implementation
Initializing the BackgroundAudioPlayer
Import the necessary classes from the voice module and instantiate the player with your desired audio configuration:
from livekit.agents import BackgroundAudioPlayer, AudioConfig, BuiltinAudioClip
background_audio = BackgroundAudioPlayer(
ambient_sound=AudioConfig(
BuiltinAudioClip.OFFICE_AMBIENCE,
volume=0.8
),
thinking_sound=[
AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING, volume=0.8),
AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING2, volume=0.7),
],
)
The ambient_sound parameter accepts a single AudioConfig, a file path string, or a list of AudioConfig objects for probability-based selection. The thinking_sound parameter triggers automatically when the agent enters the thinking state.
Configuring Ambient and Thinking Sounds
Use AudioConfig to control volume levels and selection probability:
# Probability-based ambient selection
background_audio = BackgroundAudioPlayer(
ambient_sound=[
AudioConfig(
BuiltinAudioClip.FOREST_AMBIENCE,
volume=0.6,
probability=0.4
),
AudioConfig(
BuiltinAudioClip.CITY_AMBIENCE,
volume=0.7,
probability=0.6
),
]
)
When using a list of AudioConfig objects, the player randomly selects one based on the configured probabilities when starting playback.
Starting the Player
Start the player after joining the room and initializing your AgentSession:
@server.rtc_session()
async def entrypoint(ctx: JobContext):
session = AgentSession(llm=openai.realtime.RealtimeModel())
await session.start(MyAgent(), room=ctx.room)
# Start background audio after session initialization
await background_audio.start(
room=ctx.room,
agent_session=session
)
The start() method creates the audio source, initializes the mixer, and publishes the local audio track to the LiveKit room. It also registers event listeners for agent state changes to trigger thinking sounds.
Playing Ad-Hoc Audio Clips
Play one-off sound effects at any time using the play() method:
# Play a notification sound once
handle = background_audio.play("assets/notification.ogg", loop=False)
# Wait for completion
await handle.wait_for_playout()
# Or stop early if needed
handle.stop()
The play() method returns a PlayHandle that allows you to monitor or interrupt the specific playback instance.
Cleanup and Resource Management
The BackgroundAudioPlayer automatically registers cleanup handlers. When using it with an AgentSession, cleanup occurs automatically when the session ends:
# Manual cleanup (if not using AgentSession lifecycle)
await background_audio.aclose()
The aclose() method stops all active playback, removes the audio track from the room, and releases mixer resources.
Complete Working Example
Here is the full implementation from the official example at examples/voice_agents/background_audio.py:
import asyncio
import logging
from dotenv import load_dotenv
from livekit.agents import (
Agent,
AgentServer,
AgentSession,
AudioConfig,
BackgroundAudioPlayer,
BuiltinAudioClip,
JobContext,
cli,
function_tool,
)
from livekit.plugins import openai
logging.getLogger("background-audio").setLevel(logging.INFO)
load_dotenv()
class FakeWebSearchAgent(Agent):
@function_tool
async def search_web(self, query: str) -> str:
# Simulate a long-running operation so the "thinking" sound is audible
await asyncio.sleep(5)
return "Result placeholder for " + query
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext):
session = AgentSession(llm=openai.realtime.RealtimeModel())
await session.start(FakeWebSearchAgent(), room=ctx.room)
# Background audio configuration
background_audio = BackgroundAudioPlayer(
ambient_sound=AudioConfig(BuiltinAudioClip.OFFICE_AMBIENCE, volume=0.8),
thinking_sound=[
AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING, volume=0.8),
AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING2, volume=0.7),
],
)
await background_audio.start(room=ctx.room, agent_session=session)
# Example of playing a one-off clip (uncomment to test)
# background_audio.play("assets/notification.ogg")
if __name__ == "__main__":
cli.run_app(server)
Advanced Configuration Options
Probability-Based Sound Selection
When you want variety in your ambient background, provide multiple AudioConfig objects with probability weights:
background_audio = BackgroundAudioPlayer(
ambient_sound=[
AudioConfig(
"assets/coffee_shop.ogg",
volume=0.5,
probability=0.7
),
AudioConfig(
"assets/rain.ogg",
volume=0.6,
probability=0.3
),
]
)
The player selects one configuration based on the probability distribution when initializing the ambient track.
Custom Audio Files
You can use custom OGG files instead of built-in clips:
from livekit.agents.utils import audio_frames_from_file
# Direct file path
background_audio = BackgroundAudioPlayer(
ambient_sound=AudioConfig("assets/custom_ambience.ogg", volume=0.5)
)
# Or manually create an iterator if you need processing
async def custom_audio_source():
async for frame in audio_frames_from_file("assets/processed.ogg"):
# Apply custom filters here
yield frame
Key Implementation Files
Understanding the source structure helps when debugging or extending functionality:
livekit/agents/voice/background_audio.py– Core implementation ofBackgroundAudioPlayer,AudioConfig,PlayHandle, andBuiltinAudioCliplivekit/agents/voice/agent_activity.py– Contains_interrupt_background_speeches()which stops thinking cues when the agent begins speakinglivekit/agents/voice/events.py– DefinesAgentStateChangedEventused to trigger thinking soundslivekit/agents/utils/audio.py– Providesaudio_frames_from_file()for reading OGG files into audio framesexamples/voice_agents/background_audio.py– Complete working example demonstrating ambient and thinking sounds
Common Pitfalls and Best Practices
Single Track Limitation
Each BackgroundAudioPlayer manages exactly one continuous audio track. To switch ambient sounds dynamically, stop the current ambient handle and start a new one:
# Store the handle when starting
ambient_handle = background_audio.play("assets/new_ambience.ogg", loop=True)
# Later, to switch
ambient_handle.stop()
new_handle = background_audio.play("assets/different.ogg", loop=True)
Console Mode Behavior
Background audio is automatically disabled when running in console mode (python myagent.py console) because no LiveKit room connection exists. The player logs a warning and skips publishing rather than raising an error.
Volume Clipping
The volume parameter in AudioConfig applies linear gain. Values exceeding 1.0 will cause audio clipping. Keep volumes between 0.0 and 1.0 for clean audio.
Resource Cleanup
While AgentSession automatically calls aclose() on the player when the session ends, manual cleanup is required if you manage the player lifecycle independently:
await background_audio.aclose()
Summary
- BackgroundAudioPlayer manages a single continuous audio track for ambient sounds and reactive "thinking" cues in
livekit/agents - Initialize the player with
AudioConfigobjects specifying volume and probability weights for sound selection - Call
await player.start(room=ctx.room, agent_session=session)after joining the room to begin publishing audio - The player automatically triggers thinking sounds when
AgentStateChangedEventindicates the agent is processing - Use
player.play()for ad-hoc sound effects andawait player.aclose()for cleanup - All audio mixing happens through a Rust-backed
rtc.AudioMixerfor low-latency performance
Frequently Asked Questions
How do I disable the thinking sound but keep ambient background audio?
Omit the thinking_sound parameter when initializing BackgroundAudioPlayer. Only configure ambient_sound to maintain continuous background audio without reactive cues when the agent processes requests.
Can I use MP3 files instead of OGG for background audio?
The current implementation in livekit.agents.utils.audio.audio_frames_from_file() primarily supports OGG format. Convert MP3 files to OGG before use, or implement a custom async iterator that yields rtc.AudioFrame objects from your MP3 decoder.
Why does my background audio stop when the agent starts speaking?
This is the intended behavior implemented in AgentActivity._interrupt_background_speeches(). The system automatically stops "thinking" cues and other background speeches when the agent transitions to speaking to prevent audio overlap. Ambient sounds configured with loop=True should continue playing during agent speech unless manually stopped.
How do I change the ambient sound dynamically during a session?
Store the PlayHandle returned by play() when starting your ambient sound, then call stop() on that handle when you want to switch. Create a new handle with player.play(new_sound, loop=True) to start the new ambient track. Note that only one ambient track should run at a time per player instance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →