How to Implement Text-Only Agent Mode in LiveKit Agents

TL;DR: Create an AgentSession without STT or TTS components, enable text_input=True and text_output=True in RoomOptions, and let the default text-input callback handle incoming messages by interrupting the current turn and calling generate_reply().

Text-only agent mode in the LiveKit Agents framework enables conversational AI that processes and responds exclusively via text streams, eliminating the latency and compute costs of speech-to-text and text-to-speech processing. This configuration is ideal for chat-based interfaces where clients send and receive messages on specific LiveKit topics without audio overhead.

Architecture of Text-Only Mode

Core Components

The implementation relies on three primary classes from the livekit/agents repository:

  • RoomOptions (defined in livekit/agents/voice/room_io/types.py, lines 109–121): Configures the I/O streams. Set text_input=True and text_output=True while explicitly disabling audio with audio_input=False and audio_output=False.
  • AgentSession (defined in livekit/agents/voice/agent_session.py): Manages the conversation runtime. When text input is enabled, it registers the _default_text_input_cb callback that handles incoming text messages.
  • ClientEventsHandler (defined in livekit/agents/voice/client_events.py, lines 262–279): Publishes LLM responses to the lk.transcription text stream so connected clients receive the replies.

Message Flow

The system processes text through the following pipeline:

  1. The client publishes a TextStream message on the lk.chat topic.
  2. ClientEventsHandler receives the message and invokes the registered text-input callback.
  3. The callback interrupts any ongoing generation via await sess.interrupt() and calls sess.generate_reply(user_input=text).
  4. The LLM generates a response, which AgentSession routes back to the handler.
  5. The handler publishes the reply on lk.transcription using self._room.publish_text().

Because no audio streams are configured, STT and TTS components are never instantiated.

Minimal Implementation Example

The repository provides a complete reference implementation in examples/other/text_only.py. This example demonstrates the minimal configuration required to run a text-only agent:

import logging
from dotenv import load_dotenv

from livekit.agents import (
    Agent,
    AgentServer,
    AgentSession,
    JobContext,
    cli,
    inference,
    room_io,
)

logger = logging.getLogger("text-only")
logger.setLevel(logging.INFO)
load_dotenv()


class MyAgent(Agent):
    def __init__(self) -> None:
        super().__init__(instructions="You are a helpful assistant.")


server = AgentServer()


@server.rtc_session()
async def entrypoint(ctx: JobContext):
    # Initialize session with LLM only—no audio components

    session = AgentSession(
        llm=inference.LLM("openai/gpt-4.1-mini"),
    )
    
    await session.start(
        agent=MyAgent(),
        room=ctx.room,
        room_options=room_io.RoomOptions(
            text_input=True,   # Accept TextStream on lk.chat

            text_output=True,  # Publish replies to lk.transcription

            audio_input=False,
            audio_output=False,
        ),
    )


if __name__ == "__main__":
    cli.run_app(server)

Key Configuration Details

  • Agent subclass: Defines the system prompt and behavior. No audio-related configuration is required.
  • AgentServer: Exposes HTTP and WebSocket endpoints that LiveKit clients connect to.
  • RoomOptions: The text_input and text_output booleans translate into concrete TextInputOptions objects that register the default callback.

Customizing Text Input Handling

Custom Callbacks

Replace the default _default_text_input_cb by providing your own text_input_cb in TextInputOptions:

async def my_text_cb(sess: AgentSession, ev: TextInputEvent) -> None:
    # Example: Strip non-ASCII characters before processing

    cleaned = ev.text.encode("ascii", "ignore").decode()
    await sess.interrupt()
    sess.generate_reply(user_input=cleaned)

# Pass custom handler in RoomOptions

room_options=room_io.RoomOptions(
    text_input=room_io.TextInputOptions(text_input_cb=my_text_cb),
    text_output=True,
)

Managing Interruptions

By default, new text input interrupts ongoing LLM generation. To allow the agent to complete its current turn before processing new messages, set allow_interruptions=False in the AgentSession initialization.

Integrating Tools

Tools function identically to voice mode. Pass your tool definitions to AgentSession(tools=[...]), and the LLM will invoke them as needed. Results are serialized as text and sent via the lk.transcription stream without audio processing overhead.

Running Your Text-Only Agent

Deploy the agent using the standard LiveKit CLI workflow:


# Install dependencies (includes LiveKit SDK and inference providers)

pip install livekit-agents

# Run the example server (listens on localhost:8000 by default)

python examples/other/text_only.py

Connect any LiveKit client (Web, iOS, or Android) and publish text messages on the lk.chat topic. The agent will respond on lk.transcription within milliseconds, with no audio devices required.

Summary

  • Configure RoomOptions with text_input=True and text_output=True to enable text-only mode, explicitly disabling audio streams to eliminate STT/TTS overhead.
  • The default callback (_default_text_input_cb) automatically handles incoming lk.chat messages by interrupting the current turn and triggering generate_reply().
  • Responses publish to lk.transcription via ClientEventsHandler, making them available to all connected clients as standard text streams.
  • Customization is available through TextInputOptions for pre-processing logic, interruption control, and tool integration.

Frequently Asked Questions

Do I need to configure STT or TTS for text-only mode?

No. When you set audio_input=False and audio_output=False in RoomOptions, the AgentSession skips instantiation of speech-to-text and text-to-speech components entirely. Only the LLM inference engine is required, significantly reducing startup time and runtime costs.

How does the agent handle concurrent text messages?

By default, the text-input callback calls sess.interrupt() before processing new input, which immediately cancels any ongoing LLM generation and starts a fresh reply. To queue messages instead and prevent interruptions, set allow_interruptions=False in the AgentSession configuration.

Can I use tools with text-only agents?

Yes. Tools function identically to voice mode. Pass your tool definitions to AgentSession(tools=[...]), and the LLM will invoke them as needed. The results are serialized as text and sent via the lk.transcription stream without requiring audio processing.

What LiveKit topics should clients use for text-only communication?

Clients must publish text messages on the lk.chat topic to send input to the agent. The agent publishes responses on lk.transcription. These are the default topics configured in ClientEventsHandler when text input and output are enabled in RoomOptions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →