# Building Agent Evaluation Benchmarks with Detailed Rubrics: A Complete Framework from ai-agent-book

> Learn to build robust LLM agent evaluation benchmarks with detailed rubrics. Discover a three-layer framework for reliable and auditable agent testing.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: tutorial
- Published: 2026-08-18

---

**Build reliable, auditable benchmarks for LLM agents using a three-layer architecture of test cases, system matrices, and multi-dimensional scoring rubrics with automatic hallucination vetos.**

Agent evaluation is notoriously noisy—results shift with prompt changes, model updates, or backend provider instability. The **ai-agent-book** repository by bojieli solves this with a production-grade framework for **building agent evaluation benchmarks with detailed rubrics**. This article walks through the complete architecture, from 60 curated test cases to automated backend substitution and four-dimensional rubric scoring.

---

## The Three-Layer Benchmark Architecture

The framework in `chapter7/user-memory-system-evaluation/` is organized into tightly coupled layers that together enable reproducible, multi-dimensional evaluation.

### Test-Case Layer: Stable Evaluation Surface

All experiments draw from a single set of **60 user-memory scenarios** stored in `chapter3/user-memory-evaluation/test_cases`. These cases span:

- Simple factual recall
- Multi-turn disambiguation
- Cross-session memory linking

Fixing the test set eliminates variance from case selection, making results comparable across embedding changes, model swaps, or reranker upgrades.

### System-Matrix Layer: Exhaustive Backend Combinations

The matrix is defined in [`chapter7/user-memory-system-evaluation/default_config.yaml`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/user-memory-system-evaluation/default_config.yaml). It enumerates every combination of:

| Component | Options Example |
|-----------|---------------|
| Embeddings | `bge-m3`, OpenAI `text-embedding-3-large`, Cohere |
| Rerankers | none, `bge-reranker-large`, ColBERT |
| Main LLM | Kimi, GPT-4o, Claude 3.5 Sonnet |

Each matrix cell runs the full 60-case suite, producing trajectory traces and per-cell metrics. The [`run_full.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_full.py) orchestrator handles checkpointing, retries, and parallel execution across `--workers`.

### Rubric-Scoring Layer: Deterministic Multi-Dimensional Grading

The rubric—defined in [`book/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter7.md) (lines 25–58)—enforces four graded dimensions plus a safety veto:

| Dimension | Weight | Description |
|-----------|--------|-------------|
| **Precision** | essential | Correctness of recalled facts |
| **Recall** | important | Coverage of relevant memories |
| **Reasoning** | important | Quality of explanation for selections |
| **Proactivity** | important | Appropriate initiative in ambiguous queries |
| **Hallucination Veto** | veto | Any fabricated content zeroes the score |

Weights are explicitly marked (`essential`, `important`, `veto`), making the scoring auditable and resistant to gaming.

---

## Automated Backend Substitution for Benchmark Stability

External API providers fail or rotate credentials. Rather than crash, the framework substitutes equivalent backends automatically.

Substitution rules live in [`default_config.yaml`](https://github.com/bojieli/ai-agent-book/blob/main/default_config.yaml). When SiliconFlow is unavailable, the matrix transparently routes to OpenRouter's `baai/bge-m3` equivalent. Every substitution is logged to [`results/full_matrix_backend_readiness_20260731.json`](https://github.com/bojieli/ai-agent-book/blob/main/results/full_matrix_backend_readiness_20260731.json) (see README lines 42–58), preserving experimental provenance.

This ensures **benchmarks remain runnable** without manual intervention, a critical property for longitudinal studies or CI/CD pipelines.

---

## Running the Full Evaluation Matrix

### Experiment 7-11: Complete 60-Case Run

```bash
cd chapter7/user-memory-system-evaluation
python -m pip install -r requirements.txt
cp env.example .env      # add required API keys

python run_full.py 7-11 \
    --config default_config.yaml \
    --workers 4 \
    --readiness results/full_matrix_backend_readiness.json \
    --output results/full_7_11_60_case_matrix.json

```

The orchestrator probes all backends, substitutes where needed, executes 60 cases per matrix cell, and writes a unified JSON report with hit@5, recall@5, MRR, step counts, tool-call statistics, and cost breakdowns.

### Single Configuration Testing (Experiment 7-4)

For rapid iteration on one backend combination:

```bash
python experiment.py 7-4 \
    --config default_config.yaml \
    --output results/experiment_7_4.json

```

---

## Loading and Inspecting Rubric Data

### Programmatic Rubric Access

The rubric YAML block can be extracted from the book source for custom judges:

```python
import yaml, json, pathlib

rubric_path = pathlib.Path("../book/chapter7.md")
with rubric_path.open() as f:
    for line in f:
        if line.strip().startswith("rubric:"):
            yaml_block = "\n".join(f.readlines()[0:30])
            rubric = yaml.safe_load(yaml_block)
            break

print(json.dumps(rubric, indent=2, ensure_ascii=False))

```

### Reading Per-Cell Results

```python
import json, pathlib

data = json.loads(
    pathlib.Path("results/full_7_11_60_case_matrix.json").read_text()
)

# Extract Kimi + BGE-M3 + no reranker score

score = data["Kimi"]["bge-m3"]["no-reranker"]["rubric_score"]
print("Rubric score:", score)

```

The evidence file [`results/full_7_3_structured_rubric_evidence.json`](https://github.com/bojieli/ai-agent-book/blob/main/results/full_7_3_structured_rubric_evidence.json) contains the complete four-grade judgments with hallucination vetos for all 60 cases, enabling independent verification via [`validation/verify_full_matrix_20260731.py`](https://github.com/bojieli/ai-agent-book/blob/main/validation/verify_full_matrix_20260731.py).

---

## Design Principles for Extensible Benchmarks

The ai-agent-book framework embodies four patterns transferable to any agent evaluation domain:

1. **Curated, frozen test cases** — prevents test-set contamination and enables fair comparison
2. **Explicit matrix enumeration** — makes the evaluation surface complete and inspectable
3. **Structured rubrics with veto powers** — encodes domain knowledge and safety constraints directly
4. **Backend abstraction with substitution logging** — decouples results from provider availability

These principles are discussed in the multi-modal LLM-as-a-Judge section of chapter 7, suggesting extension to vision, audio, or tool-use agents.

---

## Summary

- **Three-layer architecture**: test cases, system matrix, and rubric scoring create reproducible, auditable benchmarks
- **60 curated user-memory scenarios** provide a stable evaluation surface across all experiments
- **Four-dimensional rubric** (precision, recall, reasoning, proactivity) plus hallucination veto enforces deterministic, safety-conscious grading
- **Automatic backend substitution** with logged decisions keeps benchmarks runnable despite provider instability
- **Complete artifact chain**: from [`default_config.yaml`](https://github.com/bojieli/ai-agent-book/blob/main/default_config.yaml) through [`run_full.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_full.py) to [`full_7_11_60_case_matrix.json`](https://github.com/bojieli/ai-agent-book/blob/main/full_7_11_60_case_matrix.json) and independent [`verify_full_matrix_20260731.py`](https://github.com/bojieli/ai-agent-book/blob/main/verify_full_matrix_20260731.py) validation

---

## Frequently Asked Questions

### How does the hallucination veto work in practice?

Any detected fabrication automatically zeros the rubric score regardless of other dimensions. This is implemented in the Experiment 7-3 judge according to lines 25–58 of [`book/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book/chapter7.md). The veto appears as a boolean field in [`results/full_7_3_structured_rubric_evidence.json`](https://github.com/bojieli/ai-agent-book/blob/main/results/full_7_3_structured_rubric_evidence.json), making hallucination incidents directly auditable.

### Can I add new embeddings or models without rewriting test cases?

Yes. The matrix in [`default_config.yaml`](https://github.com/bojieli/ai-agent-book/blob/main/default_config.yaml) is designed for extension. Add new entries to the embeddings, rerankers, or models lists, and [`run_full.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_full.py) will automatically include them in the full factorial run. The 60 test cases in `chapter3/user-memory-evaluation/test_cases` remain unchanged, ensuring comparability.

### What happens if an API provider is temporarily unavailable?

The framework probes all configured backends and substitutes equivalent alternatives per rules in [`default_config.yaml`](https://github.com/bojieli/ai-agent-book/blob/main/default_config.yaml). Substitutions are logged to [`results/full_matrix_backend_readiness_20260731.json`](https://github.com/bojieli/ai-agent-book/blob/main/results/full_matrix_backend_readiness_20260731.json) with timestamps and rationale, preserving experimental reproducibility even when original providers fail.

### How do I verify that the rubric was applied consistently?

Run [`validation/verify_full_matrix_20260731.py`](https://github.com/bojieli/ai-agent-book/blob/main/validation/verify_full_matrix_20260731.py) against [`results/full_7_3_structured_rubric_evidence.json`](https://github.com/bojieli/ai-agent-book/blob/main/results/full_7_3_structured_rubric_evidence.json). This independent checker re-evaluates a sample of judgments against the rubric definition to detect scoring drift or implementation bugs.