How to Configure Turn Detection and Interruption Handling in LiveKit Agents
Configure turn detection via the turn_detection parameter in AgentSession using modes like "stt", "vad", "realtime_llm", or a custom _TurnDetector, while controlling interruptions through the allow_interruptions flag on Agent or per-speech handles.
LiveKit Agents provides granular control over conversation flow through two distinct mechanisms: turn detection (determining when a user has finished speaking) and interruption handling (whether the user can cut off the agent's response). This guide explains how to configure both using the livekit/agents repository's voice pipeline components.
Understanding Turn Detection Modes
Turn detection logic resides in livekit-agents/livekit/agents/voice/audio_recognition.py, which defines TurnDetectionMode as the central configuration type.
Built-in Literal Modes
The SDK supports four string-based modes passed to the turn_detection parameter:
"stt"– Uses Speech-to-Text partial and final results to detect end-of-utterance"vad"– Uses Voice Activity Detection (e.g., Silero) to identify speech gaps"realtime_llm"– Delegates detection to the server-side Realtime LLM API (OpenAI Realtime)"manual"– Disables automatic detection; you must callsession.commit_user_turn()programmatically
from livekit.agents import AgentSession
# Use VAD-based turn detection
session = AgentSession(
turn_detection="vad",
vad=silero.VAD.load()
)
Custom Turn Detector Objects
For advanced use cases, implement the _TurnDetector protocol (exposed by livekit-plugins-turn-detector) and pass an instance directly:
from livekit.plugins.turn_detector.multilingual import MultilingualModel
session = AgentSession(
turn_detection=MultilingualModel(), # Custom text-based detector
stt=deepgram.STT(),
llm=openai.realtime.RealtimeModel(
turn_detection=None, # Disable LLM-native detection
)
)
Automatic Mode Selection
If turn_detection is omitted, AgentSession.__init__ automatically selects the best supported mode following this priority: realtime_llm → vad → stt → manual.
Configuring Interruption Handling
Interruption logic is controlled separately from turn detection through the allow_interruptions boolean flag.
Global Agent Configuration
Set the default behavior in the Agent constructor located in livekit-agents/livekit/agents/voice/agent.py:
from livekit.agents import Agent
agent = Agent(
instructions="You are a helpful assistant.",
allow_interruptions=True # Default: users can interrupt speech
)
When allow_interruptions=False, the framework queues new user input until the current reply completes.
Per-Turn Overrides
Override the global setting for specific utterances using session.say() or session.generate_reply():
# Prevent interruption during this specific message
await session.say(
"Please hold while I connect you.",
allow_interruptions=False
)
This pattern is demonstrated in the warm_transfer.py example, where critical transfer messages must complete without interruption.
Runtime Enforcement
The SpeechHandle class in livekit-agents/livekit/agents/voice/speech_handle.py enforces interruption constraints at runtime. Attempting to disable interruptions on an already-interrupted speech handle raises a RuntimeError:
handle = await session.say("Important announcement...")
# Later, if user already interrupted:
handle.allow_interruptions = False # Raises RuntimeError
Compatibility with Realtime LLM
When using turn_detection="realtime_llm", you cannot set allow_interruptions=False. The AgentActivity class in livekit-agents/livekit/agents/voice/agent_activity.py validates this constraint:
if self._turn_detection_mode == "realtime_llm" and not allow_interruptions:
raise ValueError(
"the RealtimeModel uses a server-side turn detection, allow_interruptions cannot be False"
)
Server-side turn detection inherently manages the conversation flow, making manual interruption blocking incompatible.
Complete Implementation Example
The following example from examples/voice_agents/realtime_turn_detector.py demonstrates a production configuration using custom turn detection with interruption handling:
import logging
from dotenv import load_dotenv
from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
from livekit.plugins import deepgram, openai, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel
logger = logging.getLogger("demo")
logger.setLevel(logging.INFO)
load_dotenv()
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext):
# Load VAD for the turn detector
vad = silero.VAD.load()
# Configure session with custom turn detection and interruptions enabled
session = AgentSession(
allow_interruptions=True,
turn_detection=MultilingualModel(), # LiveKit's text-based detector
vad=vad,
stt=deepgram.STT(),
llm=openai.realtime.RealtimeModel(
voice="alloy",
turn_detection=None, # Disable LLM-native detection
input_audio_transcription=None,
),
)
await session.start(
agent=Agent(instructions="You are a helpful assistant."),
room=ctx.room,
)
def prewarm(proc):
proc.userdata["vad"] = silero.VAD.load()
server.setup_fnc = prewarm
if __name__ == "__main__":
cli.run_app(server)
This configuration:
- Uses
MultilingualModelfor turn detection instead of the Realtime LLM's built-in detection - Enables interruptions globally via
allow_interruptions=True - Disables the OpenAI Realtime model's native turn detection to avoid conflicts
Key Source Files
| File | Purpose |
|---|---|
livekit-agents/livekit/agents/voice/audio_recognition.py |
Defines TurnDetectionMode and coordinates STT/VAD/Detector logic |
livekit-agents/livekit/agents/voice/agent.py |
Agent class with allow_interruptions attribute |
livekit-agents/livekit/agents/voice/agent_session.py |
AgentSession constructor accepting turn_detection argument |
livekit-agents/livekit/agents/voice/agent_activity.py |
Validates that realtime_llm mode cannot use allow_interruptions=False |
livekit-agents/livekit/agents/voice/speech_handle.py |
Runtime enforcement of interruption constraints |
examples/voice_agents/realtime_turn_detector.py |
Production example combining custom detector with Realtime LLM |
Summary
- Turn detection determines when the user has finished speaking and is configured via the
turn_detectionparameter inAgentSessionusing modes"stt","vad","realtime_llm","manual", or a custom_TurnDetectorobject. - Interruption handling controls whether users can cut off the agent's speech through the
allow_interruptionsboolean, set globally onAgentor per-utterance viasession.say()andsession.generate_reply(). - When using
turn_detection="realtime_llm", you cannot disable interruptions (allow_interruptions=False) because the server-side detection manages conversation flow. - The
SpeechHandleclass enforces interruption constraints at runtime, raisingRuntimeErrorif you attempt to disable interruptions on already-interrupted speech.
Frequently Asked Questions
What happens if I don't specify a turn detection mode?
If you omit the turn_detection argument when constructing AgentSession, the SDK automatically selects the best available mode based on your configured components. The selection priority is: realtime_llm (if using a Realtime LLM) → vad (if VAD is available) → stt (if STT is configured) → manual (fallback).
Can I use allow_interruptions=False with OpenAI Realtime models?
No. When using turn_detection="realtime_llm" with OpenAI's Realtime API, setting allow_interruptions=False raises a ValueError in AgentActivity. Server-side turn detection inherently manages the conversation state, making manual interruption blocking incompatible. If you need to prevent interruptions, use a LiveKit-native turn detector like MultilingualModel instead.
How do I implement a custom turn detector?
Implement the _TurnDetector protocol (available through livekit-plugins-turn-detector) which requires a predict_end_of_turn coroutine and metadata attributes (model, provider). Pass your implementation instance directly to the turn_detection parameter in AgentSession. The plugin's MultilingualModel demonstrates this pattern for text-based end-of-turn prediction.
What is the difference between VAD and turn detection?
VAD (Voice Activity Detection) identifies when audio contains speech versus silence, providing a binary speech/non-speech signal. Turn detection uses VAD (or STT/LLM signals) to determine the higher-level semantic boundary of a complete user utterance. In LiveKit Agents, VAD is often a component used by turn detection, but you can use "vad" mode for simple activity-based turn detection or "stt" mode for transcript-based detection.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →