How to Implement Background Audio Handling in LiveKit Agents: A Complete Guide

Use the BackgroundAudioPlayer class from livekit.agents.voice.background_audio to manage looping ambient sounds and "thinking" cues that trigger automatically when your agent enters the thinking state.

The livekit/agents repository provides a robust background audio subsystem that lets you add immersive ambient soundscapes and reactive audio cues to your voice agents. This implementation handles audio mixing, state-driven playback, and automatic interruption when the agent needs to speak.

What Is Background Audio Handling in LiveKit Agents?

Background audio handling in LiveKit Agents refers to the continuous playback of ambient sounds (like office noise or nature sounds) alongside reactive "thinking" cues that play when the LLM is processing. The system manages a single continuous local audio track published to the LiveKit room, mixing multiple audio sources through a Rust-backed rtc.AudioMixer for low-latency performance.

The BackgroundAudioPlayer class coordinates with AgentSession to automatically interrupt background sounds when the agent begins speaking or when a new user turn arrives.

Architecture Overview

The background audio system consists of several coordinated components:

Component Role Key Methods Source File
BackgroundAudioPlayer Manages continuous audio track, ambient loops, and thinking cues start(room, agent_session), play(source, loop=False), aclose() livekit/agents/voice/background_audio.py
AudioConfig Wraps sound sources with volume and probability settings AudioConfig(source, volume=1.0, probability=1.0) livekit/agents/voice/background_audio.py
BuiltinAudioClip Enum of built-in OGG clips (office ambience, keyboard typing) BuiltinAudioClip.OFFICE_AMBIENCE.path() livekit/agents/voice/background_audio.py
PlayHandle Controls individual playback instances stop(), done() livekit/agents/voice/background_audio.py
AgentActivity Interrupts background audio on user turns _interrupt_background_speeches() livekit/agents/voice/agent_activity.py

Audio Flow:

  1. BackgroundAudioPlayer.start() creates an rtc.AudioSource and rtc.AudioMixer, publishing a LocalAudioTrack to the room
  2. Ambient sounds loop continuously through the mixer
  3. When AgentStateChangedEvent fires with new_state == "thinking", the player triggers the configured thinking sound
  4. AgentActivity._interrupt_background_speeches() stops thinking cues when the agent begins speaking

Step-by-Step Implementation

Initializing the BackgroundAudioPlayer

Import the necessary classes from the voice module and instantiate the player with your desired audio configuration:

from livekit.agents import BackgroundAudioPlayer, AudioConfig, BuiltinAudioClip

background_audio = BackgroundAudioPlayer(
    ambient_sound=AudioConfig(
        BuiltinAudioClip.OFFICE_AMBIENCE, 
        volume=0.8
    ),
    thinking_sound=[
        AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING, volume=0.8),
        AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING2, volume=0.7),
    ],
)

The ambient_sound parameter accepts a single AudioConfig, a file path string, or a list of AudioConfig objects for probability-based selection. The thinking_sound parameter triggers automatically when the agent enters the thinking state.

Configuring Ambient and Thinking Sounds

Use AudioConfig to control volume levels and selection probability:


# Probability-based ambient selection

background_audio = BackgroundAudioPlayer(
    ambient_sound=[
        AudioConfig(
            BuiltinAudioClip.FOREST_AMBIENCE, 
            volume=0.6, 
            probability=0.4
        ),
        AudioConfig(
            BuiltinAudioClip.CITY_AMBIENCE, 
            volume=0.7, 
            probability=0.6
        ),
    ]
)

When using a list of AudioConfig objects, the player randomly selects one based on the configured probabilities when starting playback.

Starting the Player

Start the player after joining the room and initializing your AgentSession:

@server.rtc_session()
async def entrypoint(ctx: JobContext):
    session = AgentSession(llm=openai.realtime.RealtimeModel())
    await session.start(MyAgent(), room=ctx.room)
    
    # Start background audio after session initialization

    await background_audio.start(
        room=ctx.room, 
        agent_session=session
    )

The start() method creates the audio source, initializes the mixer, and publishes the local audio track to the LiveKit room. It also registers event listeners for agent state changes to trigger thinking sounds.

Playing Ad-Hoc Audio Clips

Play one-off sound effects at any time using the play() method:


# Play a notification sound once

handle = background_audio.play("assets/notification.ogg", loop=False)

# Wait for completion

await handle.wait_for_playout()

# Or stop early if needed

handle.stop()

The play() method returns a PlayHandle that allows you to monitor or interrupt the specific playback instance.

Cleanup and Resource Management

The BackgroundAudioPlayer automatically registers cleanup handlers. When using it with an AgentSession, cleanup occurs automatically when the session ends:


# Manual cleanup (if not using AgentSession lifecycle)

await background_audio.aclose()

The aclose() method stops all active playback, removes the audio track from the room, and releases mixer resources.

Complete Working Example

Here is the full implementation from the official example at examples/voice_agents/background_audio.py:

import asyncio
import logging

from dotenv import load_dotenv
from livekit.agents import (
    Agent,
    AgentServer,
    AgentSession,
    AudioConfig,
    BackgroundAudioPlayer,
    BuiltinAudioClip,
    JobContext,
    cli,
    function_tool,
)
from livekit.plugins import openai

logging.getLogger("background-audio").setLevel(logging.INFO)
load_dotenv()


class FakeWebSearchAgent(Agent):
    @function_tool
    async def search_web(self, query: str) -> str:
        # Simulate a long-running operation so the "thinking" sound is audible

        await asyncio.sleep(5)
        return "Result placeholder for " + query


server = AgentServer()


@server.rtc_session()
async def entrypoint(ctx: JobContext):
    session = AgentSession(llm=openai.realtime.RealtimeModel())
    await session.start(FakeWebSearchAgent(), room=ctx.room)

    # Background audio configuration

    background_audio = BackgroundAudioPlayer(
        ambient_sound=AudioConfig(BuiltinAudioClip.OFFICE_AMBIENCE, volume=0.8),
        thinking_sound=[
            AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING, volume=0.8),
            AudioConfig(BuiltinAudioClip.KEYBOARD_TYPING2, volume=0.7),
        ],
    )
    await background_audio.start(room=ctx.room, agent_session=session)

    # Example of playing a one-off clip (uncomment to test)

    # background_audio.play("assets/notification.ogg")

    

if __name__ == "__main__":
    cli.run_app(server)

Advanced Configuration Options

Probability-Based Sound Selection

When you want variety in your ambient background, provide multiple AudioConfig objects with probability weights:

background_audio = BackgroundAudioPlayer(
    ambient_sound=[
        AudioConfig(
            "assets/coffee_shop.ogg", 
            volume=0.5, 
            probability=0.7
        ),
        AudioConfig(
            "assets/rain.ogg", 
            volume=0.6, 
            probability=0.3
        ),
    ]
)

The player selects one configuration based on the probability distribution when initializing the ambient track.

Custom Audio Files

You can use custom OGG files instead of built-in clips:

from livekit.agents.utils import audio_frames_from_file

# Direct file path

background_audio = BackgroundAudioPlayer(
    ambient_sound=AudioConfig("assets/custom_ambience.ogg", volume=0.5)
)

# Or manually create an iterator if you need processing

async def custom_audio_source():
    async for frame in audio_frames_from_file("assets/processed.ogg"):
        # Apply custom filters here

        yield frame

Key Implementation Files

Understanding the source structure helps when debugging or extending functionality:

Common Pitfalls and Best Practices

Single Track Limitation

Each BackgroundAudioPlayer manages exactly one continuous audio track. To switch ambient sounds dynamically, stop the current ambient handle and start a new one:


# Store the handle when starting

ambient_handle = background_audio.play("assets/new_ambience.ogg", loop=True)

# Later, to switch

ambient_handle.stop()
new_handle = background_audio.play("assets/different.ogg", loop=True)

Console Mode Behavior

Background audio is automatically disabled when running in console mode (python myagent.py console) because no LiveKit room connection exists. The player logs a warning and skips publishing rather than raising an error.

Volume Clipping

The volume parameter in AudioConfig applies linear gain. Values exceeding 1.0 will cause audio clipping. Keep volumes between 0.0 and 1.0 for clean audio.

Resource Cleanup

While AgentSession automatically calls aclose() on the player when the session ends, manual cleanup is required if you manage the player lifecycle independently:

await background_audio.aclose()

Summary

  • BackgroundAudioPlayer manages a single continuous audio track for ambient sounds and reactive "thinking" cues in livekit/agents
  • Initialize the player with AudioConfig objects specifying volume and probability weights for sound selection
  • Call await player.start(room=ctx.room, agent_session=session) after joining the room to begin publishing audio
  • The player automatically triggers thinking sounds when AgentStateChangedEvent indicates the agent is processing
  • Use player.play() for ad-hoc sound effects and await player.aclose() for cleanup
  • All audio mixing happens through a Rust-backed rtc.AudioMixer for low-latency performance

Frequently Asked Questions

How do I disable the thinking sound but keep ambient background audio?

Omit the thinking_sound parameter when initializing BackgroundAudioPlayer. Only configure ambient_sound to maintain continuous background audio without reactive cues when the agent processes requests.

Can I use MP3 files instead of OGG for background audio?

The current implementation in livekit.agents.utils.audio.audio_frames_from_file() primarily supports OGG format. Convert MP3 files to OGG before use, or implement a custom async iterator that yields rtc.AudioFrame objects from your MP3 decoder.

Why does my background audio stop when the agent starts speaking?

This is the intended behavior implemented in AgentActivity._interrupt_background_speeches(). The system automatically stops "thinking" cues and other background speeches when the agent transitions to speaking to prevent audio overlap. Ambient sounds configured with loop=True should continue playing during agent speech unless manually stopped.

How do I change the ambient sound dynamically during a session?

Store the PlayHandle returned by play() when starting your ambient sound, then call stop() on that handle when you want to switch. Create a new handle with player.play(new_sound, loop=True) to start the new ambient track. Note that only one ambient track should run at a time per player instance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →