How to Use the Built-In Test Framework with Judges in LiveKit Agents
LiveKit Agents provides a built-in test framework with judges that lets you write end-to-end workflow tests and automatically verify that agent responses satisfy specific intents using LLM-based evaluation.
The LiveKit Agents repository includes a purpose-built testing harness designed specifically for voice and multimodal AI agents. This built-in test framework with judges enables developers to assert not just that an agent returned a message, but that the message content aligns with a described intent, task completion, or safety criteria.
Architecture of the Built-In Test Framework
The testing infrastructure centers on two core components: the result capture mechanism that records all agent events, and the judge system that evaluates content against specified criteria.
RunResult and Event Assertions
When you execute a workflow using AgentSession.run(), the method returns a RunResult object defined in livekit/agents/voice/run_result.py. This object holds the complete event stream including messages, function calls, and handoffs.
The framework provides fluent assertion helpers that let you navigate this event stream:
RunAssert– Entry point for walking through events sequentiallynext_event()– Advances to the next event in the streamis_message(role="assistant")– Filters for chat messages and returns aChatMessageAssertcontains_function_call()– Verifies specific tool invocations occurred
These classes implement a chainable API so you can write readable test sequences like result.expect.next_event().is_message(role="assistant").
The Judge Infrastructure
The ChatMessageAssert class provides the judge() method (lines 56-84 in run_result.py), which serves as the primary interface for the built-in test framework with judges. This method sends the assistant's message content plus a textual intent description to an LLM-based evaluator.
The evaluation flow works as follows:
- The
judge()method receives anllminstance and anintentstring describing what the message should accomplish - It constructs a system prompt explaining the judge role and evaluation criteria
- The LLM calls an internal
check_intentfunction tool, returning a boolean verdict and reasoning - The method raises an assertion error if the verdict is
fail(ormaybewhen strict mode is enabled)
The underlying implementation reuses _evaluate_with_llm from livekit/agents/evals/judge.py, which provides the generic LLM-evaluation helper used by all built-in judges.
How to Use the Built-In Test Framework with Judges
Implementing tests requires executing your agent workflow, selecting specific events from the result stream, and applying judge assertions to verify intent satisfaction.
Basic Workflow Testing
The standard pattern for using the built-in test framework with judges involves three steps: running the session, navigating to the target event, and calling the judge method.
import pytest
from livekit.agents import AgentSession
from livekit.agents.llm import LLM
async def test_appointment_scheduling(sess: AgentSession, llm: LLM):
# 1️⃣ Execute the workflow
result = await sess.run(
user_input="Can I get an appointment tomorrow?"
)
# 2️⃣ Navigate to the assistant's response
# 3️⃣ Judge the intent
await (
result.expect
.next_event()
.is_message(role="assistant")
.judge(
llm,
intent=(
"must suggest one or more available appointment "
"time slots for tomorrow and only use slots that exist "
"in the calendar."
),
)
)
This pattern appears in the official test suite at tests/test_workflows.py and in the front-desk example at examples/frontdesk/test_agent.py. The judge() method returns the same ChatMessageAssert instance, enabling chained assertions if you need to verify multiple criteria.
Using Built-In Judges Directly
For higher-level evaluations such as task completion or safety checks, the framework provides reusable judge classes in livekit/agents/evals/judge.py.
from livekit.agents.evals import task_completion_judge
from livekit.agents.llm import LLM
async def test_task_completion(chat_history, llm: LLM):
# Instantiate a built-in judge
judge = task_completion_judge(llm)
# Evaluate the entire conversation context
result = await judge.evaluate(chat_history)
# Check the verdict
assert result.verdict == "pass", f"Task failed: {result.reasoning}"
print(f"Judgment: {result.verdict} - {result.reasoning}")
Built-in judges encapsulate the prompt engineering required for specific evaluation types. They use the same _evaluate_with_llm helper as the ChatMessageAssert.judge() method, ensuring consistent evaluation semantics across the framework.
Key Implementation Files
Understanding the source structure helps when extending or debugging the built-in test framework with judges:
| File | Purpose |
|---|---|
livekit/agents/voice/run_result.py |
Contains RunResult, RunAssert, ChatMessageAssert, and the judge() method implementation (lines 56-84). |
livekit/agents/evals/judge.py |
Defines JudgmentResult, Judge base class, _evaluate_with_llm helper, and built-in judges like task_completion_judge. |
tests/test_workflows.py |
Minimal test demonstrating DTMF workflow validation using .judge(). |
examples/frontdesk/test_agent.py |
Real-world example showing multiple judge assertions for a scheduling agent. |
These files demonstrate how the framework captures execution results, provides fluent assertion APIs, and integrates LLM-based evaluation for intent verification.
Summary
The built-in test framework with judges in LiveKit Agents provides a robust solution for testing AI agent workflows:
- RunResult captures complete event streams from
AgentSession.run()calls, storing messages, function calls, and handoffs inlivekit/agents/voice/run_result.py. - Fluent assertions via
RunAssertandChatMessageAssertenable readable test navigation through event sequences. - LLM-based judging through
ChatMessageAssert.judge()automatically validates that assistant messages satisfy described intents using the_evaluate_with_llmhelper inlivekit/agents/evals/judge.py. - Built-in judges like
task_completion_judgeprovide reusable evaluation patterns for common testing scenarios.
Frequently Asked Questions
How does the judge method determine if an agent response passes or fails?
The judge() method sends the assistant's message and your intent description to an LLM with a specialized system prompt. The LLM calls an internal check_intent function tool that returns a boolean verdict and reasoning. If the verdict is fail, the method raises an assertion error. If the verdict is maybe, it raises an error only when strict mode is enabled. This implementation resides in livekit/agents/voice/run_result.py lines 56-84.
Can I use the built-in test framework with judges without using LiveKit's specific LLM implementations?
Yes, the framework is designed to work with any provider-compatible LLM instance. The judge() method and built-in judges accept an llm parameter that conforms to the livekit.agents.llm.LLM interface. As long as your LLM implementation supports function calling (required for the check_intent tool), you can use OpenAI, Anthropic, or custom LLM backends with the built-in test framework with judges.
What is the difference between using ChatMessageAssert.judge() and the standalone task_completion_judge?
ChatMessageAssert.judge() is designed for inline assertions within fluent test chains, evaluating a single assistant message against a specific intent description. It integrates directly with the RunResult event stream and raises assertion errors immediately. In contrast, task_completion_judge (and other judges in livekit/agents/evals/judge.py) are standalone classes that evaluate entire ChatContext objects or conversation histories. These are better suited for holistic evaluations across multiple turns or complex task completion criteria rather than single-message assertions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →