How to Implement Transcription for Multi-User Rooms with LiveKit Agents
Use the MultiSpeakerAdapter class to wrap any diarization-enabled STT plugin, which automatically tags transcript chunks with speaker IDs and manages primary/background speaker detection via RMS energy analysis.
LiveKit Agents provides a production-ready transcription stack built around the abstract STT interface. When implementing transcription for multi-user rooms, you must identify who is speaking to maintain conversational context and prevent the agent from confusing overlapping dialogue. This guide walks you through the exact implementation using the MultiSpeakerAdapter and diarization-capable speech-to-text plugins.
Prerequisites: Diarization-Enabled STT Plugins
Before processing multiple speakers, your underlying STT service must support speaker diarization. According to the source code in livekit-agents/livekit/agents/stt/stt.py, the STT implementation must set STTCapabilities.diarization to True and populate SpeechData.speaker_id in emitted events.
Compatible plugins include:
- Speechmatics – Set
enable_diarization=Truein the constructor - Deepgram – Configure diarization parameters in the API request
The MultiSpeakerAdapter validates this capability during initialization and raises an error if the wrapped STT lacks diarization support (see validation logic at lines 19-30 in multi_speaker_adapter.py).
Understanding the MultiSpeakerAdapter Architecture
The transcription flow for multi-user rooms relies on three core components defined in livekit-agents/livekit/agents/stt/multi_speaker_adapter.py:
MultiSpeakerAdapter– Wraps the base STT and creates aMultiSpeakerAdapterWrapperstream for each recognition session.MultiSpeakerAdapterWrapper– Forwards incoming audio frames to both the underlying STT recognizer and the internal_PrimarySpeakerDetector(see the_runimplementation at lines 100-135)._PrimarySpeakerDetector– Performs RMS-based energy analysis to rank speakers, determine the primary speaker, and format or suppress transcripts according to your configuration (see lines 44-158).
When audio flows through the adapter, the detector maintains rolling RMS buffers for each speaker_id. It switches the primary speaker designation based on energy levels, then applies custom formatting rules to distinguish foreground from background speech.
Step-by-Step Implementation
1. Select a Diarization-Enabled STT Plugin
Initialize your base STT with diarization capabilities enabled. This example uses Speechmatics, but Deepgram and other providers follow similar patterns.
from livekit.plugins import speechmatics
base_stt = speechmatics.STT(
enable_diarization=True,
end_of_utterance_silence_trigger=0.3
)
2. Wrap with MultiSpeakerAdapter
Import MultiSpeakerAdapter and configure it to detect the primary speaker and format transcript output.
from livekit.agents.stt import MultiSpeakerAdapter
stt = MultiSpeakerAdapter(
stt=base_stt,
detect_primary_speaker=True, # Enable RMS-based primary detection
suppress_background_speaker=False, # Set True to drop background speech entirely
primary_format="<{speaker_id}>{text}</{speaker_id}>",
background_format="<bg>{text}</bg>",
)
3. Configure AgentSession
Pass the adapted STT instance to your AgentSession configuration. The adapter transparently handles speaker tracking without modifying your agent logic.
from livekit.agents import Agent, AgentSession, JobContext
from livekit.plugins import openai, silero
class Assistant(Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a helpful assistant.")
async def entrypoint(ctx: JobContext) -> None:
session = AgentSession(
vad=silero.VAD.load(),
llm=openai.LLM(),
tts=openai.TTS(),
stt=stt, # MultiSpeakerAdapter instance
)
await session.start(room=ctx.room, agent=Assistant())
4. Deploy the Agent
Run your agent application using the LiveKit CLI. For a complete, runnable reference implementation, see examples/voice_agents/speaker_id_multi_speaker.py in the repository.
python agent.py start
Complete Implementation Example
Here is the minimal code required to transcribe multi-user rooms with speaker identification and primary speaker formatting:
from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
from livekit.plugins import speechmatics, openai, silero
from livekit.agents.stt import MultiSpeakerAdapter
# Step 1: Initialize diarization-enabled STT
base_stt = speechmatics.STT(enable_diarization=True)
# Step 2: Wrap with MultiSpeakerAdapter for primary speaker detection
stt = MultiSpeakerAdapter(
stt=base_stt,
detect_primary_speaker=True,
suppress_background_speaker=False,
primary_format="<{speaker_id}>{text}</{speaker_id}>",
background_format="<bg>{text}</bg>",
)
class Assistant(Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a friendly assistant.")
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext) -> None:
session = AgentSession(
vad=silero.VAD.load(),
llm=openai.LLM(),
tts=openai.TTS(),
stt=stt,
)
await session.start(room=ctx.room, agent=Assistant())
if __name__ == "__main__":
cli.run_app(server)
Customizing Primary Speaker Behavior
The MultiSpeakerAdapter provides fine-grained control over how the agent perceives different speakers:
detect_primary_speaker– WhenTrue, the_PrimarySpeakerDetectorcalculates RMS energy for each speaker and designates the loudest voice as primary.suppress_background_speaker– WhenTrue, transcripts from non-primary speakers are dropped entirely rather than formatted.primary_format– A template string using{speaker_id}and{text}placeholders to markup primary speaker transcripts.background_format– A separate template for background speakers, useful for distinguishing side conversations.
Adjust these parameters based on your use case. For customer service bots, suppressing background speakers prevents interruptions, while meeting assistants might retain formatted background speech for completeness.
Summary
- Speaker diarization requires an STT plugin with
STTCapabilities.diarization=True, such as Speechmatics or Deepgram. MultiSpeakerAdapterwraps diarization-enabled STTs and is implemented inlivekit-agents/livekit/agents/stt/multi_speaker_adapter.py.- The adapter uses
_PrimarySpeakerDetectorto perform RMS-based energy analysis and identify the primary speaker in real-time. - Configure
primary_formatandbackground_formatto customize how speaker identities appear in transcripts. - Set
suppress_background_speaker=Trueto filter out non-primary speech entirely. - Reference the full example at
examples/voice_agents/speaker_id_multi_speaker.pyfor production deployment patterns.
Frequently Asked Questions
Does LiveKit Agents support real-time speaker diarization?
Yes. When you wrap a diarization-enabled STT plugin with MultiSpeakerAdapter, the system processes speaker identification in real-time as audio streams arrive. The _PrimarySpeakerDetector updates its RMS buffers continuously (see lines 44-158 in multi_speaker_adapter.py), allowing immediate primary speaker switching without batch processing delays.
Which STT plugins support speaker diarization?
The LiveKit Agents framework supports diarization through plugins that expose the speaker_id field in SpeechData events. As of the current release, Speechmatics and Deepgram officially support this capability. Verify diarization availability by checking the STTCapabilities.diarization flag on the plugin instance before wrapping it with MultiSpeakerAdapter.
How does the primary speaker detector work?
The _PrimarySpeakerDetector class calculates RMS (Root Mean Square) energy for each speaker's audio stream over a sliding window. It ranks speakers by energy level and designates the highest-energy speaker as primary. When energy levels shift—such as when a new person starts speaking—the detector switches the primary designation and applies the appropriate formatting rules defined in your MultiSpeakerAdapter configuration.
Can I suppress background speakers completely?
Yes. Set suppress_background_speaker=True when initializing MultiSpeakerAdapter. When this flag is enabled, the adapter drops all SpeechEvent objects from non-primary speakers rather than applying the background_format template. This ensures your agent processes only the dominant speaker's transcripts, which is ideal for high-noise environments or focused one-on-one interactions within group settings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →