How to Implement Transcription for Multi-User Rooms with LiveKit Agents

Use the MultiSpeakerAdapter class to wrap any diarization-enabled STT plugin, which automatically tags transcript chunks with speaker IDs and manages primary/background speaker detection via RMS energy analysis.

LiveKit Agents provides a production-ready transcription stack built around the abstract STT interface. When implementing transcription for multi-user rooms, you must identify who is speaking to maintain conversational context and prevent the agent from confusing overlapping dialogue. This guide walks you through the exact implementation using the MultiSpeakerAdapter and diarization-capable speech-to-text plugins.

Prerequisites: Diarization-Enabled STT Plugins

Before processing multiple speakers, your underlying STT service must support speaker diarization. According to the source code in livekit-agents/livekit/agents/stt/stt.py, the STT implementation must set STTCapabilities.diarization to True and populate SpeechData.speaker_id in emitted events.

Compatible plugins include:

  • Speechmatics – Set enable_diarization=True in the constructor
  • Deepgram – Configure diarization parameters in the API request

The MultiSpeakerAdapter validates this capability during initialization and raises an error if the wrapped STT lacks diarization support (see validation logic at lines 19-30 in multi_speaker_adapter.py).

Understanding the MultiSpeakerAdapter Architecture

The transcription flow for multi-user rooms relies on three core components defined in livekit-agents/livekit/agents/stt/multi_speaker_adapter.py:

  1. MultiSpeakerAdapter – Wraps the base STT and creates a MultiSpeakerAdapterWrapper stream for each recognition session.
  2. MultiSpeakerAdapterWrapper – Forwards incoming audio frames to both the underlying STT recognizer and the internal _PrimarySpeakerDetector (see the _run implementation at lines 100-135).
  3. _PrimarySpeakerDetector – Performs RMS-based energy analysis to rank speakers, determine the primary speaker, and format or suppress transcripts according to your configuration (see lines 44-158).

When audio flows through the adapter, the detector maintains rolling RMS buffers for each speaker_id. It switches the primary speaker designation based on energy levels, then applies custom formatting rules to distinguish foreground from background speech.

Step-by-Step Implementation

1. Select a Diarization-Enabled STT Plugin

Initialize your base STT with diarization capabilities enabled. This example uses Speechmatics, but Deepgram and other providers follow similar patterns.

from livekit.plugins import speechmatics

base_stt = speechmatics.STT(
    enable_diarization=True,
    end_of_utterance_silence_trigger=0.3
)

2. Wrap with MultiSpeakerAdapter

Import MultiSpeakerAdapter and configure it to detect the primary speaker and format transcript output.

from livekit.agents.stt import MultiSpeakerAdapter

stt = MultiSpeakerAdapter(
    stt=base_stt,
    detect_primary_speaker=True,          # Enable RMS-based primary detection

    suppress_background_speaker=False,    # Set True to drop background speech entirely

    primary_format="<{speaker_id}>{text}</{speaker_id}>",
    background_format="<bg>{text}</bg>",
)

3. Configure AgentSession

Pass the adapted STT instance to your AgentSession configuration. The adapter transparently handles speaker tracking without modifying your agent logic.

from livekit.agents import Agent, AgentSession, JobContext
from livekit.plugins import openai, silero

class Assistant(Agent):
    def __init__(self) -> None:
        super().__init__(instructions="You are a helpful assistant.")

async def entrypoint(ctx: JobContext) -> None:
    session = AgentSession(
        vad=silero.VAD.load(),
        llm=openai.LLM(),
        tts=openai.TTS(),
        stt=stt,  # MultiSpeakerAdapter instance

    )
    await session.start(room=ctx.room, agent=Assistant())

4. Deploy the Agent

Run your agent application using the LiveKit CLI. For a complete, runnable reference implementation, see examples/voice_agents/speaker_id_multi_speaker.py in the repository.

python agent.py start

Complete Implementation Example

Here is the minimal code required to transcribe multi-user rooms with speaker identification and primary speaker formatting:

from livekit.agents import Agent, AgentServer, AgentSession, JobContext, cli
from livekit.plugins import speechmatics, openai, silero
from livekit.agents.stt import MultiSpeakerAdapter

# Step 1: Initialize diarization-enabled STT

base_stt = speechmatics.STT(enable_diarization=True)

# Step 2: Wrap with MultiSpeakerAdapter for primary speaker detection

stt = MultiSpeakerAdapter(
    stt=base_stt,
    detect_primary_speaker=True,
    suppress_background_speaker=False,
    primary_format="<{speaker_id}>{text}</{speaker_id}>",
    background_format="<bg>{text}</bg>",
)

class Assistant(Agent):
    def __init__(self) -> None:
        super().__init__(instructions="You are a friendly assistant.")

server = AgentServer()

@server.rtc_session()
async def entrypoint(ctx: JobContext) -> None:
    session = AgentSession(
        vad=silero.VAD.load(),
        llm=openai.LLM(),
        tts=openai.TTS(),
        stt=stt,
    )
    await session.start(room=ctx.room, agent=Assistant())

if __name__ == "__main__":
    cli.run_app(server)

Customizing Primary Speaker Behavior

The MultiSpeakerAdapter provides fine-grained control over how the agent perceives different speakers:

  • detect_primary_speaker – When True, the _PrimarySpeakerDetector calculates RMS energy for each speaker and designates the loudest voice as primary.
  • suppress_background_speaker – When True, transcripts from non-primary speakers are dropped entirely rather than formatted.
  • primary_format – A template string using {speaker_id} and {text} placeholders to markup primary speaker transcripts.
  • background_format – A separate template for background speakers, useful for distinguishing side conversations.

Adjust these parameters based on your use case. For customer service bots, suppressing background speakers prevents interruptions, while meeting assistants might retain formatted background speech for completeness.

Summary

  • Speaker diarization requires an STT plugin with STTCapabilities.diarization=True, such as Speechmatics or Deepgram.
  • MultiSpeakerAdapter wraps diarization-enabled STTs and is implemented in livekit-agents/livekit/agents/stt/multi_speaker_adapter.py.
  • The adapter uses _PrimarySpeakerDetector to perform RMS-based energy analysis and identify the primary speaker in real-time.
  • Configure primary_format and background_format to customize how speaker identities appear in transcripts.
  • Set suppress_background_speaker=True to filter out non-primary speech entirely.
  • Reference the full example at examples/voice_agents/speaker_id_multi_speaker.py for production deployment patterns.

Frequently Asked Questions

Does LiveKit Agents support real-time speaker diarization?

Yes. When you wrap a diarization-enabled STT plugin with MultiSpeakerAdapter, the system processes speaker identification in real-time as audio streams arrive. The _PrimarySpeakerDetector updates its RMS buffers continuously (see lines 44-158 in multi_speaker_adapter.py), allowing immediate primary speaker switching without batch processing delays.

Which STT plugins support speaker diarization?

The LiveKit Agents framework supports diarization through plugins that expose the speaker_id field in SpeechData events. As of the current release, Speechmatics and Deepgram officially support this capability. Verify diarization availability by checking the STTCapabilities.diarization flag on the plugin instance before wrapping it with MultiSpeakerAdapter.

How does the primary speaker detector work?

The _PrimarySpeakerDetector class calculates RMS (Root Mean Square) energy for each speaker's audio stream over a sliding window. It ranks speakers by energy level and designates the highest-energy speaker as primary. When energy levels shift—such as when a new person starts speaking—the detector switches the primary designation and applies the appropriate formatting rules defined in your MultiSpeakerAdapter configuration.

Can I suppress background speakers completely?

Yes. Set suppress_background_speaker=True when initializing MultiSpeakerAdapter. When this flag is enabled, the adapter drops all SpeechEvent objects from non-primary speakers rather than applying the background_format template. This ensures your agent processes only the dominant speaker's transcripts, which is ideal for high-noise environments or focused one-on-one interactions within group settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →