# How to Use the Built-In Test Framework with Judges in LiveKit Agents

> Learn to use the LiveKit Agents built-in test framework with judges for end-to-end workflow testing. Automatically verify agent responses and LLM intents for robust development.

- Repository: [LiveKit/agents](https://github.com/livekit/agents)
- Tags: how-to-guide
- Published: 2026-03-06

---

**LiveKit Agents provides a built-in test framework with judges that lets you write end-to-end workflow tests and automatically verify that agent responses satisfy specific intents using LLM-based evaluation.**

The LiveKit Agents repository includes a purpose-built testing harness designed specifically for voice and multimodal AI agents. This built-in test framework with judges enables developers to assert not just that an agent returned a message, but that the message content aligns with a described intent, task completion, or safety criteria.

## Architecture of the Built-In Test Framework

The testing infrastructure centers on two core components: the result capture mechanism that records all agent events, and the judge system that evaluates content against specified criteria.

### RunResult and Event Assertions

When you execute a workflow using `AgentSession.run()`, the method returns a `RunResult` object defined in [`livekit/agents/voice/run_result.py`](https://github.com/livekit/agents/blob/main/livekit/agents/voice/run_result.py). This object holds the complete event stream including messages, function calls, and handoffs.

The framework provides fluent assertion helpers that let you navigate this event stream:

- **`RunAssert`** – Entry point for walking through events sequentially
- **`next_event()`** – Advances to the next event in the stream
- **`is_message(role="assistant")`** – Filters for chat messages and returns a `ChatMessageAssert`
- **`contains_function_call()`** – Verifies specific tool invocations occurred

These classes implement a chainable API so you can write readable test sequences like `result.expect.next_event().is_message(role="assistant")`.

### The Judge Infrastructure

The `ChatMessageAssert` class provides the `judge()` method (lines 56-84 in [`run_result.py`](https://github.com/livekit/agents/blob/main/run_result.py)), which serves as the primary interface for the built-in test framework with judges. This method sends the assistant's message content plus a textual intent description to an LLM-based evaluator.

The evaluation flow works as follows:

1. The `judge()` method receives an `llm` instance and an `intent` string describing what the message should accomplish
2. It constructs a system prompt explaining the judge role and evaluation criteria
3. The LLM calls an internal `check_intent` function tool, returning a boolean verdict and reasoning
4. The method raises an assertion error if the verdict is `fail` (or `maybe` when strict mode is enabled)

The underlying implementation reuses `_evaluate_with_llm` from [`livekit/agents/evals/judge.py`](https://github.com/livekit/agents/blob/main/livekit/agents/evals/judge.py), which provides the generic LLM-evaluation helper used by all built-in judges.

## How to Use the Built-In Test Framework with Judges

Implementing tests requires executing your agent workflow, selecting specific events from the result stream, and applying judge assertions to verify intent satisfaction.

### Basic Workflow Testing

The standard pattern for using the built-in test framework with judges involves three steps: running the session, navigating to the target event, and calling the judge method.

```python
import pytest
from livekit.agents import AgentSession
from livekit.agents.llm import LLM

async def test_appointment_scheduling(sess: AgentSession, llm: LLM):
    # 1️⃣ Execute the workflow

    result = await sess.run(
        user_input="Can I get an appointment tomorrow?"
    )
    
    # 2️⃣ Navigate to the assistant's response

    # 3️⃣ Judge the intent

    await (
        result.expect
              .next_event()
              .is_message(role="assistant")
              .judge(
                  llm,
                  intent=(
                      "must suggest one or more available appointment "
                      "time slots for tomorrow and only use slots that exist "
                      "in the calendar."
                  ),
              )
    )

```

This pattern appears in the official test suite at [`tests/test_workflows.py`](https://github.com/livekit/agents/blob/main/tests/test_workflows.py) and in the front-desk example at [`examples/frontdesk/test_agent.py`](https://github.com/livekit/agents/blob/main/examples/frontdesk/test_agent.py). The `judge()` method returns the same `ChatMessageAssert` instance, enabling chained assertions if you need to verify multiple criteria.

### Using Built-In Judges Directly

For higher-level evaluations such as task completion or safety checks, the framework provides reusable judge classes in [`livekit/agents/evals/judge.py`](https://github.com/livekit/agents/blob/main/livekit/agents/evals/judge.py).

```python
from livekit.agents.evals import task_completion_judge
from livekit.agents.llm import LLM

async def test_task_completion(chat_history, llm: LLM):
    # Instantiate a built-in judge

    judge = task_completion_judge(llm)
    
    # Evaluate the entire conversation context

    result = await judge.evaluate(chat_history)
    
    # Check the verdict

    assert result.verdict == "pass", f"Task failed: {result.reasoning}"
    print(f"Judgment: {result.verdict} - {result.reasoning}")

```

Built-in judges encapsulate the prompt engineering required for specific evaluation types. They use the same `_evaluate_with_llm` helper as the `ChatMessageAssert.judge()` method, ensuring consistent evaluation semantics across the framework.

## Key Implementation Files

Understanding the source structure helps when extending or debugging the built-in test framework with judges:

| File | Purpose |
|------|---------|
| [`livekit/agents/voice/run_result.py`](https://github.com/livekit/agents/blob/main/livekit/agents/voice/run_result.py) | Contains `RunResult`, `RunAssert`, `ChatMessageAssert`, and the `judge()` method implementation (lines 56-84). |
| [`livekit/agents/evals/judge.py`](https://github.com/livekit/agents/blob/main/livekit/agents/evals/judge.py) | Defines `JudgmentResult`, `Judge` base class, `_evaluate_with_llm` helper, and built-in judges like `task_completion_judge`. |
| [`tests/test_workflows.py`](https://github.com/livekit/agents/blob/main/tests/test_workflows.py) | Minimal test demonstrating DTMF workflow validation using `.judge()`. |
| [`examples/frontdesk/test_agent.py`](https://github.com/livekit/agents/blob/main/examples/frontdesk/test_agent.py) | Real-world example showing multiple judge assertions for a scheduling agent. |

These files demonstrate how the framework captures execution results, provides fluent assertion APIs, and integrates LLM-based evaluation for intent verification.

## Summary

The built-in test framework with judges in LiveKit Agents provides a robust solution for testing AI agent workflows:

- **RunResult** captures complete event streams from `AgentSession.run()` calls, storing messages, function calls, and handoffs in [`livekit/agents/voice/run_result.py`](https://github.com/livekit/agents/blob/main/livekit/agents/voice/run_result.py).
- **Fluent assertions** via `RunAssert` and `ChatMessageAssert` enable readable test navigation through event sequences.
- **LLM-based judging** through `ChatMessageAssert.judge()` automatically validates that assistant messages satisfy described intents using the `_evaluate_with_llm` helper in [`livekit/agents/evals/judge.py`](https://github.com/livekit/agents/blob/main/livekit/agents/evals/judge.py).
- **Built-in judges** like `task_completion_judge` provide reusable evaluation patterns for common testing scenarios.

## Frequently Asked Questions

### How does the judge method determine if an agent response passes or fails?

The `judge()` method sends the assistant's message and your intent description to an LLM with a specialized system prompt. The LLM calls an internal `check_intent` function tool that returns a boolean verdict and reasoning. If the verdict is `fail`, the method raises an assertion error. If the verdict is `maybe`, it raises an error only when strict mode is enabled. This implementation resides in [`livekit/agents/voice/run_result.py`](https://github.com/livekit/agents/blob/main/livekit/agents/voice/run_result.py) lines 56-84.

### Can I use the built-in test framework with judges without using LiveKit's specific LLM implementations?

Yes, the framework is designed to work with any provider-compatible `LLM` instance. The `judge()` method and built-in judges accept an `llm` parameter that conforms to the `livekit.agents.llm.LLM` interface. As long as your LLM implementation supports function calling (required for the `check_intent` tool), you can use OpenAI, Anthropic, or custom LLM backends with the built-in test framework with judges.

### What is the difference between using `ChatMessageAssert.judge()` and the standalone `task_completion_judge`?

`ChatMessageAssert.judge()` is designed for inline assertions within fluent test chains, evaluating a single assistant message against a specific intent description. It integrates directly with the `RunResult` event stream and raises assertion errors immediately. In contrast, `task_completion_judge` (and other judges in [`livekit/agents/evals/judge.py`](https://github.com/livekit/agents/blob/main/livekit/agents/evals/judge.py)) are standalone classes that evaluate entire `ChatContext` objects or conversation histories. These are better suited for holistic evaluations across multiple turns or complex task completion criteria rather than single-message assertions.