# How to Build Evaluation-Driven Model Selection for Agent Systems

> Build evaluation-driven model selection for agent systems. Use reproducible benchmarks and LLM-as-a-Judge to score agent performance. Optimize models with statistical evidence while managing latency and cost.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Evaluation-driven model selection replaces reputation-based heuristics with statistical evidence by sandboxing agents in reproducible benchmarks, scoring trajectories via LLM-as-a-Judge, and promoting models only when performance gains exceed confidence intervals while respecting latency and cost constraints.**

The bojieli/ai-agent-book repository provides a comprehensive framework for implementing evaluation-driven model selection in production AI agent systems. Detailed in [`book/chapter6.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter6.md) and implemented across the Chapter 7 codebase, this approach treats model selection as a controlled experiment rather than a guessing game, utilizing executable outcome checks and statistical significance testing to validate every deployment decision.

## Core Architecture of Evaluation-Driven Model Selection

According to the bojieli/ai-agent-book source code, the framework organizes evaluation-driven model selection into five distinct layers that guarantee reproducible, statistically valid decisions.

### Sandbox Evaluation Environments

The foundation rests on two paradigms—tool-calling and human-computer interaction—that generate executable outcome checks. As emphasized in [`slides/lesson-04.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-04.md), these environments ensure that metrics remain comparable across runs by validating that model selection follows evaluation data rather than reputation.

### Curated Benchmark Datasets

The repository integrates standardized benchmarks including GAIA, OSWorld, SWE-bench, and Tau2-Bench, each containing structured rubrics that define success criteria and cost metrics. These datasets reside in paths like `chapter7/GAIA/` and are accessible via the clone script referenced in the README, providing ground-truth answers for rigorous agent testing.

### LLM-as-a-Judge Scoring Backend

Automated evaluation consumes agent trajectories and rubrics to return numeric scores, implemented in the test suite [`tests/test_ch7_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch7_evaluate_multilingual.py). This component eliminates manual grading bottlenecks while maintaining consistency across multilingual and multi-domain tasks.

### Statistical Significance and Ablations

After batch runs, the framework computes confidence intervals and performs paired t-tests to generate ablation reports, as detailed in [`slides/lesson-25.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-25.md). This layer surfaces which configuration changes—whether prompt tweaks or tool additions—produce genuine improvements versus statistical noise.

### Closed-Loop Feedback Integration

The pipeline feeds results back into the codebase via feature flags and prompt-sensitivity checks, enabling continuous improvement driven by evaluation data rather than intuition.

## Step-by-Step Implementation Pipeline

To implement evaluation-driven model selection in your own systems, follow the executable workflow defined in [`chapter7/evaluation/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/evaluation/main.py) and the runner modules.

### Define Target Metrics and Constraints

Begin by selecting measurable outcomes such as keyword-recall, cost-efficiency, or task success rate. Establish latency and cost thresholds upfront, as these constraints form the selection boundary alongside raw performance scores.

### Execute Controlled Benchmark Runs

Use the `run_evaluation()` function from [`chapter7/evaluation/runner.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/evaluation/runner.py) to execute agents against curated datasets like GAIA. The runner records full trajectories and stores results in JSON format within `chapter7/evaluation/results/`, ensuring every decision is traceable to raw data.

### Automate Trajectory Scoring

Invoke the LLM-as-a-Judge to evaluate each trajectory against the structured rubric, producing numeric scores that populate the JSON reports. This automated scoring mechanism, demonstrated in [`tests/test_ch7_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch7_evaluate_multilingual.py), supports batch processing across multiple languages and domains.

### Apply Statistical Selection Criteria

Aggregate scores across random seeds, compute 95% confidence intervals, and perform paired t-tests to validate significance. According to [`slides/lesson-24.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-24.md), promote a model only when its expected improvement exceeds statistical noise and fits within the predefined latency and cost envelope.

## Runnable Code Examples

The following snippets demonstrate how to invoke the evaluation pipeline using the actual implementation from bojieli/ai-agent-book.

```python

# example.py – run a multilingual evaluation batch

from chapter7.evaluation.runner import run_evaluation  # ↗ source

from chapter7.evaluation.datasets import GAIA   # ↗ source

# 1️⃣ Load a multilingual agent (any class that implements `__call__(prompt)`)

my_agent = MyMultilingualAgent(model_name="qwen2.5-7b")

# 2️⃣ Choose a dataset (GAIA is a multilingual knowledge‑base benchmark)

dataset = GAIA.load(split="test")

# 3️⃣ Execute the evaluation; a dict of scores is returned

report = run_evaluation(agent=my_agent, dataset=dataset)

print("Average score:", report["mean"])
print("95 % CI:", report["ci_95"])

```

```bash

# Bash – run the full batch and generate a PDF report

uv run python -m chapter7.evaluation.main \
    --dataset GAIA \
    --metric keyword-recall \
    --output results/gaia_report.pdf

```

The Python example imports from `chapter7.evaluation.runner` and `chapter7.evaluation.datasets`, while the CLI entry point is defined in [`chapter7/evaluation/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/evaluation/main.py). Both methods generate statistical reports that feed directly into the model selection decision process.

## Summary

- **Evaluation-driven model selection** requires sandboxed environments, curated benchmarks, automated scoring, and statistical validation to replace intuition with data.
- The bojieli/ai-agent-book implementation centers on [`chapter7/evaluation/runner.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/evaluation/runner.py) for execution and [`tests/test_ch7_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch7_evaluate_multilingual.py) for validation logic.
- Key configuration and philosophy reside in [`slides/lesson-04.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-04.md) (reputation vs. evaluation), [`slides/lesson-24.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-24.md) (selection criteria), and [`slides/lesson-25.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-25.md) (statistical methods).
- Selection decisions must satisfy both performance metrics (exceeding confidence intervals) and operational constraints (latency/cost), with all results feeding back via closed-loop integration.

## Frequently Asked Questions

### What is evaluation-driven model selection?

Evaluation-driven model selection is a methodology that chooses AI agents based on quantitative benchmark performance rather than reputation or intuition. According to the bojieli/ai-agent-book framework, this approach requires running candidates through standardized tests, applying statistical significance checks, and respecting latency and cost constraints before deployment.

### How does the LLM-as-a-Judge component work?

The LLM-as-a-Judge consumes agent trajectories and predefined rubrics to return numeric scores automatically. This implementation lives in [`tests/test_ch7_evaluate_multilingual.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_ch7_evaluate_multilingual.py) and supports multilingual evaluation, eliminating manual grading while maintaining consistency across diverse task domains.

### Which statistical methods verify model improvements?

The framework computes confidence intervals and performs paired t-tests on batch results to distinguish genuine improvements from random noise. As detailed in [`slides/lesson-25.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-25.md), these methods generate ablation reports that identify which specific changes—such as prompt adjustments or tool additions—statistically improve agent performance.

### Where are the benchmark datasets located?

Curated datasets including GAIA, OSWorld, and SWE-bench are referenced in the Chapter 7 clone script and stored in paths like `chapter7/GAIA/`. Each dataset includes structured rubrics and ground-truth answers, providing the standardized inputs necessary for reproducible evaluation-driven model selection.