How to Implement Text-Only Agent Mode in LiveKit Agents
TL;DR: Create an AgentSession without STT or TTS components, enable text_input=True and text_output=True in RoomOptions, and let the default text-input callback handle incoming messages by interrupting the current turn and calling generate_reply().
Text-only agent mode in the LiveKit Agents framework enables conversational AI that processes and responds exclusively via text streams, eliminating the latency and compute costs of speech-to-text and text-to-speech processing. This configuration is ideal for chat-based interfaces where clients send and receive messages on specific LiveKit topics without audio overhead.
Architecture of Text-Only Mode
Core Components
The implementation relies on three primary classes from the livekit/agents repository:
RoomOptions(defined inlivekit/agents/voice/room_io/types.py, lines 109–121): Configures the I/O streams. Settext_input=Trueandtext_output=Truewhile explicitly disabling audio withaudio_input=Falseandaudio_output=False.AgentSession(defined inlivekit/agents/voice/agent_session.py): Manages the conversation runtime. When text input is enabled, it registers the_default_text_input_cbcallback that handles incoming text messages.ClientEventsHandler(defined inlivekit/agents/voice/client_events.py, lines 262–279): Publishes LLM responses to thelk.transcriptiontext stream so connected clients receive the replies.
Message Flow
The system processes text through the following pipeline:
- The client publishes a
TextStreammessage on thelk.chattopic. ClientEventsHandlerreceives the message and invokes the registered text-input callback.- The callback interrupts any ongoing generation via
await sess.interrupt()and callssess.generate_reply(user_input=text). - The LLM generates a response, which
AgentSessionroutes back to the handler. - The handler publishes the reply on
lk.transcriptionusingself._room.publish_text().
Because no audio streams are configured, STT and TTS components are never instantiated.
Minimal Implementation Example
The repository provides a complete reference implementation in examples/other/text_only.py. This example demonstrates the minimal configuration required to run a text-only agent:
import logging
from dotenv import load_dotenv
from livekit.agents import (
Agent,
AgentServer,
AgentSession,
JobContext,
cli,
inference,
room_io,
)
logger = logging.getLogger("text-only")
logger.setLevel(logging.INFO)
load_dotenv()
class MyAgent(Agent):
def __init__(self) -> None:
super().__init__(instructions="You are a helpful assistant.")
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: JobContext):
# Initialize session with LLM only—no audio components
session = AgentSession(
llm=inference.LLM("openai/gpt-4.1-mini"),
)
await session.start(
agent=MyAgent(),
room=ctx.room,
room_options=room_io.RoomOptions(
text_input=True, # Accept TextStream on lk.chat
text_output=True, # Publish replies to lk.transcription
audio_input=False,
audio_output=False,
),
)
if __name__ == "__main__":
cli.run_app(server)
Key Configuration Details
Agentsubclass: Defines the system prompt and behavior. No audio-related configuration is required.AgentServer: Exposes HTTP and WebSocket endpoints that LiveKit clients connect to.RoomOptions: Thetext_inputandtext_outputbooleans translate into concreteTextInputOptionsobjects that register the default callback.
Customizing Text Input Handling
Custom Callbacks
Replace the default _default_text_input_cb by providing your own text_input_cb in TextInputOptions:
async def my_text_cb(sess: AgentSession, ev: TextInputEvent) -> None:
# Example: Strip non-ASCII characters before processing
cleaned = ev.text.encode("ascii", "ignore").decode()
await sess.interrupt()
sess.generate_reply(user_input=cleaned)
# Pass custom handler in RoomOptions
room_options=room_io.RoomOptions(
text_input=room_io.TextInputOptions(text_input_cb=my_text_cb),
text_output=True,
)
Managing Interruptions
By default, new text input interrupts ongoing LLM generation. To allow the agent to complete its current turn before processing new messages, set allow_interruptions=False in the AgentSession initialization.
Integrating Tools
Tools function identically to voice mode. Pass your tool definitions to AgentSession(tools=[...]), and the LLM will invoke them as needed. Results are serialized as text and sent via the lk.transcription stream without audio processing overhead.
Running Your Text-Only Agent
Deploy the agent using the standard LiveKit CLI workflow:
# Install dependencies (includes LiveKit SDK and inference providers)
pip install livekit-agents
# Run the example server (listens on localhost:8000 by default)
python examples/other/text_only.py
Connect any LiveKit client (Web, iOS, or Android) and publish text messages on the lk.chat topic. The agent will respond on lk.transcription within milliseconds, with no audio devices required.
Summary
- Configure
RoomOptionswithtext_input=Trueandtext_output=Trueto enable text-only mode, explicitly disabling audio streams to eliminate STT/TTS overhead. - The default callback (
_default_text_input_cb) automatically handles incominglk.chatmessages by interrupting the current turn and triggeringgenerate_reply(). - Responses publish to
lk.transcriptionviaClientEventsHandler, making them available to all connected clients as standard text streams. - Customization is available through
TextInputOptionsfor pre-processing logic, interruption control, and tool integration.
Frequently Asked Questions
Do I need to configure STT or TTS for text-only mode?
No. When you set audio_input=False and audio_output=False in RoomOptions, the AgentSession skips instantiation of speech-to-text and text-to-speech components entirely. Only the LLM inference engine is required, significantly reducing startup time and runtime costs.
How does the agent handle concurrent text messages?
By default, the text-input callback calls sess.interrupt() before processing new input, which immediately cancels any ongoing LLM generation and starts a fresh reply. To queue messages instead and prevent interruptions, set allow_interruptions=False in the AgentSession configuration.
Can I use tools with text-only agents?
Yes. Tools function identically to voice mode. Pass your tool definitions to AgentSession(tools=[...]), and the LLM will invoke them as needed. The results are serialized as text and sent via the lk.transcription stream without requiring audio processing.
What LiveKit topics should clients use for text-only communication?
Clients must publish text messages on the lk.chat topic to send input to the agent. The agent publishes responses on lk.transcription. These are the default topics configured in ClientEventsHandler when text input and output are enabled in RoomOptions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →