# How to Configure VAD in LiveKit Agents: A Complete Guide

> Easily configure Voice Activity Detection VAD in LiveKit Agents. Load plugins like Silero, set thresholds, and update options at runtime for seamless integration. Get the complete guide now.

- Repository: [LiveKit/agents](https://github.com/livekit/agents)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Configure VAD in LiveKit Agents by loading a VAD plugin (such as Silero) with custom thresholds via the `load()` class method, passing the instance to the `Agent` constructor's `vad` parameter, and optionally updating detection parameters at runtime using `update_options()`.**

LiveKit Agents provides a pluggable architecture for Voice Activity Detection (VAD) that enables precise control over speech detection in real-time voice applications. Whether you're building conversational AI agents or voice-controlled interfaces, knowing how to configure VAD in LiveKit Agents allows you to optimize turn-taking, handle interruptions, and reduce latency by detecting exactly when users start and stop speaking.

## Understanding the VAD Architecture in LiveKit Agents

The VAD system in LiveKit Agents follows a modular design that separates the core abstraction from concrete implementations.

### Core Abstractions and Event Model

The foundational VAD interface is defined in [`livekit-agents/livekit/agents/vad.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/vad.py). This module declares the abstract `VAD` class and the `VADStream` interface, along with the event model consisting of `VADEvent` and `VADEventType`. These abstractions allow the `Agent` class to consume VAD events without depending on specific implementations.

When audio flows through the system, the `AudioRecognition` component (located in [`livekit-agents/livekit/agents/voice/audio_recognition.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/voice/audio_recognition.py)) creates a dedicated async channel (`self._vad_ch`) and launches a background task (`self._vad_task`) that streams audio frames into `vad.stream()`. Detected events—`START_OF_SPEECH`, `INFERENCE_DONE`, and `END_OF_SPEECH`—are forwarded to user-defined hooks.

### The Silero ONNX Implementation

The reference VAD implementation uses the Silero ONNX-based detector, located in [`livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py`](https://github.com/livekit/agents/blob/main/livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py). This concrete implementation provides the `load()` class method that constructs an ONNX inference session and creates a `_VADOptions` instance to store tunable thresholds including `min_speech_duration`, `min_silence_duration`, and `activation_threshold`.

## How to Configure VAD in LiveKit Agents: Step-by-Step

Configuring VAD requires three main steps: loading a plugin with custom parameters, passing the instance to your agent, and optionally handling VAD events.

### Step 1: Load a VAD Plugin with Custom Thresholds

Import the Silero plugin and call `load()` with your desired detection parameters. This method builds the ONNX session and returns a configured `VAD` instance.

```python
from livekit.plugins import silero

my_vad = silero.VAD.load(
    min_speech_duration=0.3,      # seconds of speech before a turn starts

    min_silence_duration=0.6,     # seconds of silence needed to end a turn

    activation_threshold=0.55,    # probability above which a frame is speech

    sample_rate=16000,
)

```

The `activation_threshold` controls sensitivity—lower values detect quieter speech but may increase false positives, while higher values require louder speech. The `min_speech_duration` and `min_silence_duration` parameters prevent spurious turn changes due to brief pauses or noise.

### Step 2: Pass the VAD Instance to Your Agent

Pass the loaded VAD instance to the `Agent` constructor via the `vad` parameter. This enables VAD-driven turn detection within the `AudioRecognition` pipeline.

```python
from livekit.agents import Agent

agent = Agent(
    vad=my_vad,                     # enable VAD-based turn detection

    turn_detection="vad",           # explicit mode (optional—inferred automatically)

)

```

Setting `vad` to a `VAD` instance enables the VAD stream; passing `None` or omitting the parameter disables VAD processing. The `turn_detection` parameter is optional since the system infers the mode automatically when a VAD instance is present.

### Step 3: Handle VAD Events in Your Agent Subclass

Subclass `Agent` to implement hooks that respond to speech detection events. The `AudioRecognition` component forwards `START_OF_SPEECH`, `INFERENCE_DONE`, and `END_OF_SPEECH` events to these methods.

```python
class EchoAgent(Agent):
    async def on_start_of_speech(self, ev):
        print(f"Speech started at {ev.timestamp}")

    async def on_vad_inference_done(self, ev):
        print(f"VAD inference completed")

    async def on_end_of_speech(self, ev):
        print(f"Speech ended after {ev.speech_duration}s")

# Run the agent

if __name__ == "__main__":
    EchoAgent(vad=my_vad).run_console()

```

These hooks enable real-time reactions such as visual indicators when users speak, logging for analytics, or custom interruption handling logic.

## Runtime VAD Configuration and Metrics

LiveKit Agents allows dynamic adjustment of VAD parameters and provides observability into detection performance.

### Updating VAD Options Dynamically

After initialization, modify detection thresholds without rebuilding the ONNX session using `update_options()`. This updates the internal `_VADOptions` used by all active `VADStream` instances.

```python

# Adjust for a quieter environment

my_vad.update_options(
    activation_threshold=0.45,
    min_silence_duration=0.4,
)

```

This capability is essential for adaptive systems that adjust sensitivity based on ambient noise levels or user preferences during a session.

### Monitoring VAD Performance Metrics

Every `VADStream` emits `VADMetrics` periodically through the `metrics_collected` event. These metrics include idle time, inference duration, and event counts, enabling performance monitoring and optimization.

According to the source code in [`livekit-agents/livekit/agents/vad.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/vad.py), these metrics help identify latency bottlenecks in the inference pipeline and optimize resource utilization in production deployments.

## Summary

- **Load a VAD plugin** using the `load()` class method (e.g., `silero.VAD.load()`) with custom thresholds for speech duration, silence duration, and activation probability.
- **Enable VAD in your Agent** by passing the loaded instance to the `vad` parameter of the `Agent` constructor, which wires it into the `AudioRecognition` pipeline.
- **Handle speech events** by subclassing `Agent` and implementing `on_start_of_speech()`, `on_end_of_speech()`, and `on_vad_inference_done()` hooks for real-time reactions.
- **Adjust dynamically** using `update_options()` to modify thresholds without restarting the inference session, and monitor performance via `VADMetrics` emitted by each `VADStream`.

## Frequently Asked Questions

### What is the default VAD implementation in LiveKit Agents?

The reference implementation is the **Silero ONNX-based VAD**, located in [`livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py`](https://github.com/livekit/agents/blob/main/livekit-plugins/livekit-plugins-silero/livekit/plugins/silero/vad.py). This plugin provides the `load()` method and supports runtime configuration via `update_options()`. While the core abstraction in [`livekit-agents/livekit/agents/vad.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/vad.py) allows for alternative implementations, Silero is the primary supported plugin.

### How do I disable VAD in LiveKit Agents?

To disable VAD processing, pass `None` to the `vad` parameter of the `Agent` constructor, or simply omit the parameter entirely. According to the source in [`livekit-agents/livekit/agents/voice/agent.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/voice/agent.py), the type annotation is `NotGivenOr[vad.VAD | None]`, where `None` explicitly disables VAD-driven turn detection in the `AudioRecognition` pipeline.

### Can I use a custom VAD implementation instead of Silero?

Yes. The VAD architecture is pluggable. You can implement the abstract `VAD` class and `VADStream` interface defined in [`livekit-agents/livekit/agents/vad.py`](https://github.com/livekit/agents/blob/main/livekit-agents/livekit/agents/vad.py). Your implementation must provide the `stream()` method that returns a `VADStream` and emits `VADEvent` objects with types `START_OF_SPEECH`, `INFERENCE_DONE`, and `END_OF_SPEECH`. Once implemented, pass your custom instance to the `Agent` constructor just like the Silero plugin.

### What are the optimal threshold values for noisy environments?

For noisy environments, lower the `activation_threshold` (e.g., from 0.55 to 0.45) to detect quieter speech, but increase `min_speech_duration` (e.g., to 0.5 seconds) to filter out brief noise spikes. Conversely, reduce `min_silence_duration` (e.g., to 0.3 seconds) to detect turns faster when users pause briefly. Use `update_options()` to adjust these dynamically based on ambient noise levels detected during the session.