How to Configure Turn Detection and Interruption Handling in LiveKit Agents

Configure turn detection via the turn_detection parameter in AgentSession using modes like "stt", "vad", "realtime_llm", or a custom _TurnDetector, while controlling interruptions through the allow_interruptions flag on Agent or per-speech handles.

LiveKit Agents provides granular control over conversation flow through two distinct mechanisms: turn detection (determining when a user has finished speaking) and interruption handling (whether the user can cut off the agent's response). This guide explains how to configure both using the livekit/agents repository's voice pipeline components.

Understanding Turn Detection Modes

Turn detection logic resides in livekit-agents/livekit/agents/voice/audio_recognition.py, which defines TurnDetectionMode as the central configuration type.

Built-in Literal Modes

The SDK supports four string-based modes passed to the turn_detection parameter:

  • "stt" – Uses Speech-to-Text partial and final results to detect end-of-utterance
  • "vad" – Uses Voice Activity Detection (e.g., Silero) to identify speech gaps
  • "realtime_llm" – Delegates detection to the server-side Realtime LLM API (OpenAI Realtime)
  • "manual" – Disables automatic detection; you must call session.commit_user_turn() programmatically
from livekit.agents import AgentSession

# Use VAD-based turn detection

session = AgentSession(
    turn_detection="vad",
    vad=silero.VAD.load()
)

Custom Turn Detector Objects

For advanced use cases, implement the _TurnDetector protocol (exposed by livekit-plugins-turn-detector) and pass an instance directly:

from livekit.plugins.turn_detector.multilingual import MultilingualModel

session = AgentSession(
    turn_detection=MultilingualModel(),  # Custom text-based detector

    stt=deepgram.STT(),
    llm=openai.realtime.RealtimeModel(
        turn_detection=None,  # Disable LLM-native detection

    )
)

Automatic Mode Selection

If turn_detection is omitted, AgentSession.__init__ automatically selects the best supported mode following this priority: realtime_llm → vad → stt → manual.

Configuring Interruption Handling

Interruption logic is controlled separately from turn detection through the allow_interruptions boolean flag.

Global Agent Configuration

Set the default behavior in the Agent constructor located in livekit-agents/livekit/agents/voice/agent.py:

from livekit.agents import Agent

agent = Agent(
    instructions="You are a helpful assistant.",
    allow_interruptions=True  # Default: users can interrupt speech

)

When allow_interruptions=False, the framework queues new user input until the current reply completes.

Per-Turn Overrides

Override the global setting for specific utterances using session.say() or session.generate_reply():


# Prevent interruption during this specific message

await session.say(
    "Please hold while I connect you.",
    allow_interruptions=False
)

This pattern is demonstrated in the warm_transfer.py example, where critical transfer messages must complete without interruption.

Runtime Enforcement

The SpeechHandle class in livekit-agents/livekit/agents/voice/speech_handle.py enforces interruption constraints at runtime. Attempting to disable interruptions on an already-interrupted speech handle raises a RuntimeError:

handle = await session.say("Important announcement...")

# Later, if user already interrupted:

handle.allow_interruptions = False  # Raises RuntimeError

Compatibility with Realtime LLM

When using turn_detection="realtime_llm", you cannot set allow_interruptions=False. The AgentActivity class in livekit-agents/livekit/agents/voice/agent_activity.py validates this constraint:

if self._turn_detection_mode == "realtime_llm" and not allow_interruptions:
    raise ValueError(
        "the RealtimeModel uses a server-side turn detection, allow_interruptions cannot be False"
    )

Server-side turn detection inherently manages the conversation flow, making manual interruption blocking incompatible.

Complete Implementation Example

The following example from examples/voice_agents/realtime_turn_detector.py demonstrates a production configuration using custom turn detection with interruption handling:

import logging
from dotenv import load_dotenv
from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
from livekit.plugins import deepgram, openai, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel

logger = logging.getLogger("demo")
logger.setLevel(logging.INFO)

load_dotenv()
server = AgentServer()

@server.rtc_session()
async def entrypoint(ctx: JobContext):
    # Load VAD for the turn detector

    vad = silero.VAD.load()

    # Configure session with custom turn detection and interruptions enabled

    session = AgentSession(
        allow_interruptions=True,
        turn_detection=MultilingualModel(),  # LiveKit's text-based detector

        vad=vad,
        stt=deepgram.STT(),
        llm=openai.realtime.RealtimeModel(
            voice="alloy",
            turn_detection=None,  # Disable LLM-native detection

            input_audio_transcription=None,
        ),
    )

    await session.start(
        agent=Agent(instructions="You are a helpful assistant."),
        room=ctx.room,
    )

def prewarm(proc):
    proc.userdata["vad"] = silero.VAD.load()

server.setup_fnc = prewarm

if __name__ == "__main__":
    cli.run_app(server)

This configuration:

  • Uses MultilingualModel for turn detection instead of the Realtime LLM's built-in detection
  • Enables interruptions globally via allow_interruptions=True
  • Disables the OpenAI Realtime model's native turn detection to avoid conflicts

Key Source Files

File Purpose
livekit-agents/livekit/agents/voice/audio_recognition.py Defines TurnDetectionMode and coordinates STT/VAD/Detector logic
livekit-agents/livekit/agents/voice/agent.py Agent class with allow_interruptions attribute
livekit-agents/livekit/agents/voice/agent_session.py AgentSession constructor accepting turn_detection argument
livekit-agents/livekit/agents/voice/agent_activity.py Validates that realtime_llm mode cannot use allow_interruptions=False
livekit-agents/livekit/agents/voice/speech_handle.py Runtime enforcement of interruption constraints
examples/voice_agents/realtime_turn_detector.py Production example combining custom detector with Realtime LLM

Summary

  • Turn detection determines when the user has finished speaking and is configured via the turn_detection parameter in AgentSession using modes "stt", "vad", "realtime_llm", "manual", or a custom _TurnDetector object.
  • Interruption handling controls whether users can cut off the agent's speech through the allow_interruptions boolean, set globally on Agent or per-utterance via session.say() and session.generate_reply().
  • When using turn_detection="realtime_llm", you cannot disable interruptions (allow_interruptions=False) because the server-side detection manages conversation flow.
  • The SpeechHandle class enforces interruption constraints at runtime, raising RuntimeError if you attempt to disable interruptions on already-interrupted speech.

Frequently Asked Questions

What happens if I don't specify a turn detection mode?

If you omit the turn_detection argument when constructing AgentSession, the SDK automatically selects the best available mode based on your configured components. The selection priority is: realtime_llm (if using a Realtime LLM) → vad (if VAD is available) → stt (if STT is configured) → manual (fallback).

Can I use allow_interruptions=False with OpenAI Realtime models?

No. When using turn_detection="realtime_llm" with OpenAI's Realtime API, setting allow_interruptions=False raises a ValueError in AgentActivity. Server-side turn detection inherently manages the conversation state, making manual interruption blocking incompatible. If you need to prevent interruptions, use a LiveKit-native turn detector like MultilingualModel instead.

How do I implement a custom turn detector?

Implement the _TurnDetector protocol (available through livekit-plugins-turn-detector) which requires a predict_end_of_turn coroutine and metadata attributes (model, provider). Pass your implementation instance directly to the turn_detection parameter in AgentSession. The plugin's MultilingualModel demonstrates this pattern for text-based end-of-turn prediction.

What is the difference between VAD and turn detection?

VAD (Voice Activity Detection) identifies when audio contains speech versus silence, providing a binary speech/non-speech signal. Turn detection uses VAD (or STT/LLM signals) to determine the higher-level semantic boundary of a complete user utterance. In LiveKit Agents, VAD is often a component used by turn detection, but you can use "vad" mode for simple activity-based turn detection or "stt" mode for transcript-based detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →