How to Use the Built-In Test Framework with Judges in LiveKit Agents

LiveKit Agents provides a built-in test framework with judges that lets you write end-to-end workflow tests and automatically verify that agent responses satisfy specific intents using LLM-based evaluation.

The LiveKit Agents repository includes a purpose-built testing harness designed specifically for voice and multimodal AI agents. This built-in test framework with judges enables developers to assert not just that an agent returned a message, but that the message content aligns with a described intent, task completion, or safety criteria.

Architecture of the Built-In Test Framework

The testing infrastructure centers on two core components: the result capture mechanism that records all agent events, and the judge system that evaluates content against specified criteria.

RunResult and Event Assertions

When you execute a workflow using AgentSession.run(), the method returns a RunResult object defined in livekit/agents/voice/run_result.py. This object holds the complete event stream including messages, function calls, and handoffs.

The framework provides fluent assertion helpers that let you navigate this event stream:

  • RunAssert – Entry point for walking through events sequentially
  • next_event() – Advances to the next event in the stream
  • is_message(role="assistant") – Filters for chat messages and returns a ChatMessageAssert
  • contains_function_call() – Verifies specific tool invocations occurred

These classes implement a chainable API so you can write readable test sequences like result.expect.next_event().is_message(role="assistant").

The Judge Infrastructure

The ChatMessageAssert class provides the judge() method (lines 56-84 in run_result.py), which serves as the primary interface for the built-in test framework with judges. This method sends the assistant's message content plus a textual intent description to an LLM-based evaluator.

The evaluation flow works as follows:

  1. The judge() method receives an llm instance and an intent string describing what the message should accomplish
  2. It constructs a system prompt explaining the judge role and evaluation criteria
  3. The LLM calls an internal check_intent function tool, returning a boolean verdict and reasoning
  4. The method raises an assertion error if the verdict is fail (or maybe when strict mode is enabled)

The underlying implementation reuses _evaluate_with_llm from livekit/agents/evals/judge.py, which provides the generic LLM-evaluation helper used by all built-in judges.

How to Use the Built-In Test Framework with Judges

Implementing tests requires executing your agent workflow, selecting specific events from the result stream, and applying judge assertions to verify intent satisfaction.

Basic Workflow Testing

The standard pattern for using the built-in test framework with judges involves three steps: running the session, navigating to the target event, and calling the judge method.

import pytest
from livekit.agents import AgentSession
from livekit.agents.llm import LLM

async def test_appointment_scheduling(sess: AgentSession, llm: LLM):
    # 1️⃣ Execute the workflow

    result = await sess.run(
        user_input="Can I get an appointment tomorrow?"
    )
    
    # 2️⃣ Navigate to the assistant's response

    # 3️⃣ Judge the intent

    await (
        result.expect
              .next_event()
              .is_message(role="assistant")
              .judge(
                  llm,
                  intent=(
                      "must suggest one or more available appointment "
                      "time slots for tomorrow and only use slots that exist "
                      "in the calendar."
                  ),
              )
    )

This pattern appears in the official test suite at tests/test_workflows.py and in the front-desk example at examples/frontdesk/test_agent.py. The judge() method returns the same ChatMessageAssert instance, enabling chained assertions if you need to verify multiple criteria.

Using Built-In Judges Directly

For higher-level evaluations such as task completion or safety checks, the framework provides reusable judge classes in livekit/agents/evals/judge.py.

from livekit.agents.evals import task_completion_judge
from livekit.agents.llm import LLM

async def test_task_completion(chat_history, llm: LLM):
    # Instantiate a built-in judge

    judge = task_completion_judge(llm)
    
    # Evaluate the entire conversation context

    result = await judge.evaluate(chat_history)
    
    # Check the verdict

    assert result.verdict == "pass", f"Task failed: {result.reasoning}"
    print(f"Judgment: {result.verdict} - {result.reasoning}")

Built-in judges encapsulate the prompt engineering required for specific evaluation types. They use the same _evaluate_with_llm helper as the ChatMessageAssert.judge() method, ensuring consistent evaluation semantics across the framework.

Key Implementation Files

Understanding the source structure helps when extending or debugging the built-in test framework with judges:

File Purpose
livekit/agents/voice/run_result.py Contains RunResult, RunAssert, ChatMessageAssert, and the judge() method implementation (lines 56-84).
livekit/agents/evals/judge.py Defines JudgmentResult, Judge base class, _evaluate_with_llm helper, and built-in judges like task_completion_judge.
tests/test_workflows.py Minimal test demonstrating DTMF workflow validation using .judge().
examples/frontdesk/test_agent.py Real-world example showing multiple judge assertions for a scheduling agent.

These files demonstrate how the framework captures execution results, provides fluent assertion APIs, and integrates LLM-based evaluation for intent verification.

Summary

The built-in test framework with judges in LiveKit Agents provides a robust solution for testing AI agent workflows:

  • RunResult captures complete event streams from AgentSession.run() calls, storing messages, function calls, and handoffs in livekit/agents/voice/run_result.py.
  • Fluent assertions via RunAssert and ChatMessageAssert enable readable test navigation through event sequences.
  • LLM-based judging through ChatMessageAssert.judge() automatically validates that assistant messages satisfy described intents using the _evaluate_with_llm helper in livekit/agents/evals/judge.py.
  • Built-in judges like task_completion_judge provide reusable evaluation patterns for common testing scenarios.

Frequently Asked Questions

How does the judge method determine if an agent response passes or fails?

The judge() method sends the assistant's message and your intent description to an LLM with a specialized system prompt. The LLM calls an internal check_intent function tool that returns a boolean verdict and reasoning. If the verdict is fail, the method raises an assertion error. If the verdict is maybe, it raises an error only when strict mode is enabled. This implementation resides in livekit/agents/voice/run_result.py lines 56-84.

Can I use the built-in test framework with judges without using LiveKit's specific LLM implementations?

Yes, the framework is designed to work with any provider-compatible LLM instance. The judge() method and built-in judges accept an llm parameter that conforms to the livekit.agents.llm.LLM interface. As long as your LLM implementation supports function calling (required for the check_intent tool), you can use OpenAI, Anthropic, or custom LLM backends with the built-in test framework with judges.

What is the difference between using ChatMessageAssert.judge() and the standalone task_completion_judge?

ChatMessageAssert.judge() is designed for inline assertions within fluent test chains, evaluating a single assistant message against a specific intent description. It integrates directly with the RunResult event stream and raises assertion errors immediately. In contrast, task_completion_judge (and other judges in livekit/agents/evals/judge.py) are standalone classes that evaluate entire ChatContext objects or conversation histories. These are better suited for holistic evaluations across multiple turns or complex task completion criteria rather than single-message assertions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →