How to Implement Semantic Turn Detection in LiveKit Agents: 3 Methods Explained
To implement semantic turn detection in LiveKit Agents, pass a TurnDetection configuration or MultilingualModel instance to the AgentSession constructor while disabling the provider's native VAD to avoid conflicts.
LiveKit Agents is an open-source framework for building real-time voice AI applications. Semantic turn detection enables your voice agents to identify when a user has finished speaking by analyzing transcript content rather than relying solely on audio energy levels, resulting in more natural conversations and lower latency.
What Is Semantic Turn Detection?
Traditional voice activity detection (VAD) relies on audio energy thresholds to determine when someone stops speaking. Semantic turn detection (also called semantic VAD) uses language models to predict end-of-utterance (EOU) based on the conversational context. This approach works reliably even with whispered speech, background noise, or across multiple languages.
According to the source code in livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/utils.py (lines 28-33), LiveKit provides a DEFAULT_TURN_DETECTION configuration that uses type="semantic_vad" with medium eagerness and interrupt capabilities enabled.
Architecture and Key Components
The semantic turn detection system in LiveKit Agents consists of these core components:
| Component | Source File | Purpose |
|---|---|---|
to_turn_detection |
livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/utils.py (lines 27-35) |
Normalizes user configuration into OpenAI RealtimeAudioInputTurnDetection objects |
MultilingualModel |
livekit-plugins/livekit-plugins-turn-detector/livekit/plugins/turn_detector/multilingual.py |
Provides language-aware end-of-turn prediction with optional remote inference |
AgentSession |
livekit-agents/livekit/agents/voice/speech_handle.py |
Runtime orchestrator that wires VAD, STT, and LLM components |
RealtimeModel |
livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/realtime_model.py |
OpenAI Realtime API wrapper that consumes turn detection configs |
Method 1: Using the Multilingual Turn Detector Plugin
The MultilingualModel class provides semantic end-of-utterance detection across many languages without relying on OpenAI's built-in semantic VAD. This approach is ideal when you need language-specific turn detection or want to use a custom inference service.
Key features:
- Implements
predict_end_of_turnfor context-aware EOU prediction - Supports remote inference via the
LIVEKIT_REMOTE_EOT_URLenvironment variable - Integrates with
AgentSessionthrough theturn_detectionparameter
from livekit.plugins.turn_detector.multilingual import MultilingualModel
session = AgentSession(
turn_detection=MultilingualModel(), # Uses LiveKit's semantic detector
vad=silero.VAD.load(),
stt=deepgram.STT(),
llm=openai.realtime.RealtimeModel(
voice="alloy",
turn_detection=None, # Disable OpenAI's native VAD
),
)
As shown in examples/voice_agents/realtime_turn_detector.py, this configuration disables the built-in OpenAI VAD (turn_detection=None) and injects LiveKit's multilingual model instead.
Method 2: Native OpenAI Realtime Semantic VAD
For OpenAI Realtime API users, you can configure semantic VAD directly through the TurnDetection model. The to_turn_detection utility in utils.py (lines 90-121) converts these configurations into SemanticVad instances.
from openai.types.beta.realtime.session import TurnDetection
custom_vad = TurnDetection(
type="semantic_vad",
create_response=True,
eagerness="medium",
interrupt_response=True,
)
session = AgentSession(
turn_detection=custom_vad,
llm=openai.realtime.RealtimeModel(turn_detection=None),
)
When type="semantic_vad" is specified, the Realtime API performs end-of-utterance detection in the text domain rather than using raw audio energy.
Method 3: Default Semantic VAD Configuration
If you provide no turn_detection argument to AgentSession, the system automatically applies DEFAULT_TURN_DETECTION. As defined in livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/utils.py (lines 28-33), this defaults to:
type="semantic_vad"create_response=Trueeagerness="medium"interrupt_response=True
To explicitly use defaults:
from livekit.plugins.openai.realtime.utils import DEFAULT_TURN_DETECTION
session = AgentSession(
turn_detection=DEFAULT_TURN_DETECTION,
llm=openai.realtime.RealtimeModel(turn_detection=None),
)
Complete Implementation Example
The following example from examples/voice_agents/realtime_turn_detector.py demonstrates a production-ready setup using the multilingual plugin:
import logging
from dotenv import load_dotenv
from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
from livekit.plugins import openai, deepgram, silero
from livekit.plugins.turn_detector.multilingual import MultilingualModel
logger = logging.getLogger("semantic-vad-demo")
load_dotenv()
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext):
session = AgentSession(
turn_detection=MultilingualModel(),
vad=ctx.proc.userdata["vad"],
stt=deepgram.STT(),
llm=openai.realtime.RealtimeModel(
voice="alloy",
turn_detection=None, # Critical: disable OpenAI's VAD
input_audio_transcription=None,
),
)
await session.start(
agent=Agent(instructions="You are a helpful assistant."),
room=ctx.room
)
def prewarm(proc):
proc.userdata["vad"] = silero.VAD.load()
server.setup_fnc = prewarm
if __name__ == "__main__":
cli.run_app(server)
Key implementation details:
turn_detection=MultilingualModel()injects the semantic detectoropenai.realtime.RealtimeModel(turn_detection=None)prevents double VAD processing- The separate
vadparameter handles low-latency speech activity detection while the semantic detector handles turn completion
Configuration Parameters
When using TurnDetection objects, these parameters control behavior:
eagerness: Controls how aggressively the model predicts end-of-turn. Options are"low","medium", or"high". Higher values reduce latency but may cause interruptions.create_response: WhenTrue, automatically triggers LLM response generation upon turn detection.interrupt_response: WhenTrue, allows new user speech to interrupt ongoing LLM responses.
For remote multilingual inference, configure the endpoint via the LIVEKIT_REMOTE_EOT_URL environment variable as referenced in multilingual.py.
Summary
- Semantic turn detection uses language models rather than audio energy to detect when users finish speaking
- Implement via
MultilingualModel()for cross-language support orTurnDetection(type="semantic_vad")for OpenAI Realtime integration - Always disable the provider's native VAD (
turn_detection=None) when using LiveKit's semantic detection to prevent conflicts - Configure
eagerness,create_response, andinterrupt_responseto balance latency against accuracy - Core logic resides in
livekit-plugins/livekit-plugins-openai/livekit/plugins/openai/realtime/utils.pyandlivekit-plugins/livekit-plugins-turn-detector/livekit/plugins/turn_detector/multilingual.py
Frequently Asked Questions
How does semantic turn detection differ from server VAD?
Server VAD relies on audio energy thresholds and silence detection, while semantic turn detection analyzes the transcript content using language models to predict when a user has completed their thought. According to the source code, semantic VAD works even when audio energy is low (whispers) or in noisy environments, and it supports multilingual detection through the MultilingualModel class.
Can I use semantic turn detection with providers other than OpenAI?
Yes. While the TurnDetection object targets OpenAI's Realtime API structure, the MultilingualModel plugin in livekit-plugins-turn-detector provides provider-agnostic semantic detection. You can use MultilingualModel() with any STT and LLM combination by passing it to AgentSession(turn_detection=...).
What happens if I enable both OpenAI's VAD and LiveKit's turn detection?
This creates a conflict where both systems compete to detect turns. The examples/voice_agents/realtime_turn_detector.py file explicitly sets turn_detection=None on the RealtimeModel when using MultilingualModel(), ensuring only LiveKit's detector manages the conversation flow.
How do I adjust the sensitivity of turn detection?
For OpenAI's semantic VAD, set the eagerness parameter to "low", "medium", or "high" in the TurnDetection configuration (lines 90-121 of utils.py). For the multilingual model, you can adjust thresholds in the predict_end_of_turn method or configure the remote inference service via LIVEKIT_REMOTE_EOT_URL for custom probability cutoffs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →