# How LLM-as-a-Judge Works for Automated Agent Evaluation in ai-agent-book

> Discover how LLM-as-a-Judge automates agent evaluation using language models for scoring open-ended behaviors. Learn this effective technique for AI agent assessment.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-06

---

**LLM-as-a-Judge is an evaluation pattern that uses a language model to score open-ended, language-centric dimensions of an agent's behavior while keeping deterministic safety checks separate.**

The *ai-agent-book* repository by bojieli implements this pattern to automatically evaluate AI agent trajectories—complete execution traces of an agent's reasoning and actions. This article breaks down the architecture, interfaces, and implementation details based on the source code in `chapter8/trajectory-verifier/`.

## The Core Problem: Why Use LLM-as-a-Judge?

Traditional agent evaluation relies on exact match or rule-based metrics. These work for deterministic outcomes but fail for nuanced qualities like instruction following, tone appropriateness, or flexible compliance.

The repository solves this by splitting evaluation into two layers: **deterministic verifiers** for safety-critical checks and **LLM judges** for subjective, language-based assessment. This separation lets the system automate scoring while flagging edge cases for human review.

## The QualityJudge Protocol: The Plug-in Interface

All judges implement the `QualityJudge` protocol defined in [`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)【[`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)†L28-L33】. This single-method interface ensures interchangeable implementations:

```python
from typing import Protocol, Iterable

class QualityJudge(Protocol):
    def evaluate(self, trajectory: dict) -> Iterable[DimensionResult]:
        """Score a trajectory across quality dimensions."""
        ...

```

Any component—deterministic heuristic or full LLM—can slot into the evaluation pipeline through this interface.

## Deterministic Verification Layers

Before invoking any LLM, the system runs fast, rule-based checks. These never call external models and operate directly on the trajectory structure.

`ResultVerifier` and `ProcessVerifier` in [`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)【[`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)†L75-L119】 validate:

- **Policy compliance** — did the agent violate hard constraints?
- **Privacy leaks** — did it expose sensitive data?
- **Factual grounding** — are claims supported by retrieved evidence?
- **Promise-action consistency** — did the agent do what it said it would?

These verifiers return `DimensionResult` objects with pass/fail verdicts. They form the safety backbone that runs regardless of LLM availability or cost constraints.

## LLM-Based Quality Judges

For dimensions requiring language understanding, the repository provides two implementations with identical interfaces.

### HeuristicQualityJudge: Deterministic Stand-in

`HeuristicQualityJudge` synthesizes scores from pre-computed `quality_facts` in the trajectory【[`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)†L15-L58】. It requires no external API calls and produces reproducible results—ideal for unit tests and offline demonstrations.

### OpenAIQualityJudge: Real LLM Evaluation

`OpenAIQualityJudge` in [`llm_judge.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_judge.py)【[`llm_judge.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_judge.py)†L31-L73】 makes actual LLM calls through an `EvidenceChatClient`. The implementation handles:

- **Prompt engineering** — structured requests for specific rubric dimensions
- **JSON parsing** with robust error recovery
- **Null value coercion** to safe defaults【[`llm_judge.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_judge.py)†L83-L115】

The judge requests two primary dimensions: **expression_quality** and **compliant_flexibility**. Each dimension returns verdict, score, confidence level, and evidence citations【[`llm_judge.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_judge.py)†L62-L71】.

```python
from chapter8.trajectory_verifier.llm_judge import OpenAIQualityJudge

# Initialize with model selection via OpenRouter

judge = OpenAIQualityJudge(model="gpt-4o-mini")

# Or use a more capable model for critical evaluations

judge = OpenAIQualityJudge(model="gpt-4o")

```

## The Complete Evaluation Flow

`TrajectoryVerifier.evaluate` orchestrates the full pipeline【[`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)†L68-L75】:

1. Run deterministic verifiers (`ResultVerifier`, `ProcessVerifier`)
2. Delegate to configured `QualityJudge` for language-centric dimensions
3. Aggregate scores across all dimensions
4. Identify critical failures and low-confidence verdicts
5. Set human-review flags when needed【[`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py)†L78-L95】

```python
from chapter8.trajectory_verifier.verifier import TrajectoryVerifier
from chapter8.trajectory_verifier.llm_judge import OpenAIQualityJudge

# Load trajectory with messages, tool_calls, environment states

trajectory = {
    "messages": [...],
    "tool_calls": [...],
    "final_state": {...}
}

# Configure verifier with LLM judge and confidence threshold

verifier = TrajectoryVerifier(
    quality_judge=OpenAIQualityJudge(model="gpt-4o-mini"),
    review_confidence=0.75  # Flag for human review below this threshold

)

# Execute evaluation

report = verifier.evaluate(trajectory)

# Check results

print(f"Overall score: {report['overall_score']}")
print(f"Dimensions scored: {len(report['dimensions'])}")
print(f"Human review required: {report['review']['required']}")

```

## Integration in Experiments

Experiment 8-1 demonstrates real usage. The driver script instantiates `TrajectoryVerifier` with `OpenAIQualityJudge` and processes each trajectory through the unified interface【[`run_experiment_8_1.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_experiment_8_1.py)†L127-L129】:

```python

# From run_experiment_8_1.py

verifier = TrajectoryVerifier(
    quality_judge=OpenAIQualityJudge(model=args.judge_model),
    review_confidence=args.review_threshold
)
report = verifier.evaluate(trajectory)
results.append(report)

```

The same code path works with `HeuristicQualityJudge` when LLM access is unavailable—demonstrating the value of the protocol-based design.

## Error Handling and Robustness

The LLM judge includes defensive parsing for real-world API unpredictability. Tests in [`test_llm_judge_null_fields.py`](https://github.com/bojieli/ai-agent-book/blob/main/test_llm_judge_null_fields.py) verify behavior when the LLM returns:

- Explicit `null` values for optional fields
- Missing keys in the JSON response
- Malformed JSON requiring regex extraction

The parser coerces all cases to safe defaults rather than failing the entire evaluation.

## Key Files in the Repository

| Path | Purpose |
|------|---------|
| [`chapter8/trajectory-verifier/verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/trajectory-verifier/verifier.py) | Core interfaces and orchestration |
| [`chapter8/trajectory-verifier/llm_judge.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/trajectory-verifier/llm_judge.py) | OpenAI-based judge implementation |
| [`chapter8/trajectory-verifier/run_experiment_8_1.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/trajectory-verifier/run_experiment_8_1.py) | Experiment driver with LLM judge |
| [`chapter8/trajectory-verifier/test_llm_judge_null_fields.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/trajectory-verifier/test_llm_judge_null_fields.py) | Robustness test suite |
| [`chapter8/trajectory-verifier/demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/trajectory-verifier/demo.py) | Interactive demonstration script |

## Summary

- **LLM-as-a-Judge** in `ai-agent-book` uses a protocol-based architecture with `QualityJudge` as the central abstraction
- Deterministic verifiers handle safety-critical checks without LLM dependency
- `OpenAIQualityJudge` provides real LLM evaluation with structured rubric outputs and robust error handling
- `HeuristicQualityJudge` enables fast, reproducible testing without API calls
- The `TrajectoryVerifier` orchestrates both layers and flags uncertain cases for human review
- All implementations are swappable through the shared protocol, supporting experimentation and production deployment

## Frequently Asked Questions

### What is LLM-as-a-Judge in agent evaluation?

LLM-as-a-Judge is an architectural pattern where a language model scores subjective, open-ended qualities of an agent's behavior—such as instruction following, tone, or flexible compliance—that resist rule-based measurement. In `ai-agent-book`, this pattern is implemented through the `QualityJudge` protocol, which cleanly separates deterministic safety checks from nuanced language assessment.

### How does the repository handle LLM API failures or malformed responses?

`OpenAIQualityJudge` implements defensive parsing with multiple fallback strategies. When the LLM returns malformed JSON, missing fields, or explicit `null` values, the parser attempts regex extraction and coerces values to safe defaults rather than raising exceptions. This ensures evaluation pipelines remain robust against real-world API unpredictability.

### Can I use this evaluation system without OpenAI API access?

Yes. The `HeuristicQualityJudge` provides a deterministic implementation that synthesizes scores from pre-computed `quality_facts` in the trajectory. Because both judges implement the same `QualityJudge` protocol, you can swap them without changing calling code—enabling unit testing, offline development, and cost-sensitive deployments.

### What triggers a human review flag in the evaluation?

The `TrajectoryVerifier` sets `review["required"] = True` when any dimension reports a critical failure, when the LLM judge returns a low-confidence verdict below the configured `review_confidence` threshold, or when uncertain classifications appear. This creates a safety net ensuring automated scoring does not silently approve problematic agent behaviors.