How to Configure VAD in LiveKit Agents: A Complete Guide

Configure VAD in LiveKit Agents by loading a VAD plugin (such as Silero) with custom thresholds via the load() class method, passing the instance to the Agent constructor's vad parameter, and optionally updating detection parameters at runtime using update_options().

LiveKit Agents provides a pluggable architecture for Voice Activity Detection (VAD) that enables precise control over speech detection in real-time voice applications. Whether you're building conversational AI agents or voice-controlled interfaces, knowing how to configure VAD in LiveKit Agents allows you to optimize turn-taking, handle interruptions, and reduce latency by detecting exactly when users start and stop speaking.

Understanding the VAD Architecture in LiveKit Agents

The VAD system in LiveKit Agents follows a modular design that separates the core abstraction from concrete implementations.

Core Abstractions and Event Model

The foundational VAD interface is defined in livekit-agents/livekit/agents/vad.py. This module declares the abstract VAD class and the VADStream interface, along with the event model consisting of VADEvent and VADEventType. These abstractions allow the Agent class to consume VAD events without depending on specific implementations.

When audio flows through the system, the AudioRecognition component (located in livekit-agents/livekit/agents/voice/audio_recognition.py) creates a dedicated async channel (self._vad_ch) and launches a background task (self._vad_task) that streams audio frames into vad.stream(). Detected events—START_OF_SPEECH, INFERENCE_DONE, and END_OF_SPEECH—are forwarded to user-defined hooks.

The Silero ONNX Implementation

The reference VAD implementation uses the Silero ONNX-based detector, located in livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py. This concrete implementation provides the load() class method that constructs an ONNX inference session and creates a _VADOptions instance to store tunable thresholds including min_speech_duration, min_silence_duration, and activation_threshold.

How to Configure VAD in LiveKit Agents: Step-by-Step

Configuring VAD requires three main steps: loading a plugin with custom parameters, passing the instance to your agent, and optionally handling VAD events.

Step 1: Load a VAD Plugin with Custom Thresholds

Import the Silero plugin and call load() with your desired detection parameters. This method builds the ONNX session and returns a configured VAD instance.

from livekit.plugins import silero

my_vad = silero.VAD.load(
    min_speech_duration=0.3,      # seconds of speech before a turn starts

    min_silence_duration=0.6,     # seconds of silence needed to end a turn

    activation_threshold=0.55,    # probability above which a frame is speech

    sample_rate=16000,
)

The activation_threshold controls sensitivity—lower values detect quieter speech but may increase false positives, while higher values require louder speech. The min_speech_duration and min_silence_duration parameters prevent spurious turn changes due to brief pauses or noise.

Step 2: Pass the VAD Instance to Your Agent

Pass the loaded VAD instance to the Agent constructor via the vad parameter. This enables VAD-driven turn detection within the AudioRecognition pipeline.

from livekit.agents import Agent

agent = Agent(
    vad=my_vad,                     # enable VAD-based turn detection

    turn_detection="vad",           # explicit mode (optional—inferred automatically)

)

Setting vad to a VAD instance enables the VAD stream; passing None or omitting the parameter disables VAD processing. The turn_detection parameter is optional since the system infers the mode automatically when a VAD instance is present.

Step 3: Handle VAD Events in Your Agent Subclass

Subclass Agent to implement hooks that respond to speech detection events. The AudioRecognition component forwards START_OF_SPEECH, INFERENCE_DONE, and END_OF_SPEECH events to these methods.

class EchoAgent(Agent):
    async def on_start_of_speech(self, ev):
        print(f"Speech started at {ev.timestamp}")

    async def on_vad_inference_done(self, ev):
        print(f"VAD inference completed")

    async def on_end_of_speech(self, ev):
        print(f"Speech ended after {ev.speech_duration}s")

# Run the agent

if __name__ == "__main__":
    EchoAgent(vad=my_vad).run_console()

These hooks enable real-time reactions such as visual indicators when users speak, logging for analytics, or custom interruption handling logic.

Runtime VAD Configuration and Metrics

LiveKit Agents allows dynamic adjustment of VAD parameters and provides observability into detection performance.

Updating VAD Options Dynamically

After initialization, modify detection thresholds without rebuilding the ONNX session using update_options(). This updates the internal _VADOptions used by all active VADStream instances.


# Adjust for a quieter environment

my_vad.update_options(
    activation_threshold=0.45,
    min_silence_duration=0.4,
)

This capability is essential for adaptive systems that adjust sensitivity based on ambient noise levels or user preferences during a session.

Monitoring VAD Performance Metrics

Every VADStream emits VADMetrics periodically through the metrics_collected event. These metrics include idle time, inference duration, and event counts, enabling performance monitoring and optimization.

According to the source code in livekit-agents/livekit/agents/vad.py, these metrics help identify latency bottlenecks in the inference pipeline and optimize resource utilization in production deployments.

Summary

  • Load a VAD plugin using the load() class method (e.g., silero.VAD.load()) with custom thresholds for speech duration, silence duration, and activation probability.
  • Enable VAD in your Agent by passing the loaded instance to the vad parameter of the Agent constructor, which wires it into the AudioRecognition pipeline.
  • Handle speech events by subclassing Agent and implementing on_start_of_speech(), on_end_of_speech(), and on_vad_inference_done() hooks for real-time reactions.
  • Adjust dynamically using update_options() to modify thresholds without restarting the inference session, and monitor performance via VADMetrics emitted by each VADStream.

Frequently Asked Questions

What is the default VAD implementation in LiveKit Agents?

The reference implementation is the Silero ONNX-based VAD, located in livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py. This plugin provides the load() method and supports runtime configuration via update_options(). While the core abstraction in livekit-agents/livekit/agents/vad.py allows for alternative implementations, Silero is the primary supported plugin.

How do I disable VAD in LiveKit Agents?

To disable VAD processing, pass None to the vad parameter of the Agent constructor, or simply omit the parameter entirely. According to the source in livekit-agents/livekit/agents/voice/agent.py, the type annotation is NotGivenOr[vad.VAD | None], where None explicitly disables VAD-driven turn detection in the AudioRecognition pipeline.

Can I use a custom VAD implementation instead of Silero?

Yes. The VAD architecture is pluggable. You can implement the abstract VAD class and VADStream interface defined in livekit-agents/livekit/agents/vad.py. Your implementation must provide the stream() method that returns a VADStream and emits VADEvent objects with types START_OF_SPEECH, INFERENCE_DONE, and END_OF_SPEECH. Once implemented, pass your custom instance to the Agent constructor just like the Silero plugin.

What are the optimal threshold values for noisy environments?

For noisy environments, lower the activation_threshold (e.g., from 0.55 to 0.45) to detect quieter speech, but increase min_speech_duration (e.g., to 0.5 seconds) to filter out brief noise spikes. Conversely, reduce min_silence_duration (e.g., to 0.3 seconds) to detect turns faster when users pause briefly. Use update_options() to adjust these dynamically based on ambient noise levels detected during the session.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →