Building Agent Evaluation Benchmarks with Detailed Rubrics: A Complete Framework from ai-agent-book

Build reliable, auditable benchmarks for LLM agents using a three-layer architecture of test cases, system matrices, and multi-dimensional scoring rubrics with automatic hallucination vetos.

Agent evaluation is notoriously noisy—results shift with prompt changes, model updates, or backend provider instability. The ai-agent-book repository by bojieli solves this with a production-grade framework for building agent evaluation benchmarks with detailed rubrics. This article walks through the complete architecture, from 60 curated test cases to automated backend substitution and four-dimensional rubric scoring.


The Three-Layer Benchmark Architecture

The framework in chapter7/user-memory-system-evaluation/ is organized into tightly coupled layers that together enable reproducible, multi-dimensional evaluation.

Test-Case Layer: Stable Evaluation Surface

All experiments draw from a single set of 60 user-memory scenarios stored in chapter3/user-memory-evaluation/test_cases. These cases span:

  • Simple factual recall
  • Multi-turn disambiguation
  • Cross-session memory linking

Fixing the test set eliminates variance from case selection, making results comparable across embedding changes, model swaps, or reranker upgrades.

System-Matrix Layer: Exhaustive Backend Combinations

The matrix is defined in chapter7/user-memory-system-evaluation/default_config.yaml. It enumerates every combination of:

Component Options Example
Embeddings bge-m3, OpenAI text-embedding-3-large, Cohere
Rerankers none, bge-reranker-large, ColBERT
Main LLM Kimi, GPT-4o, Claude 3.5 Sonnet

Each matrix cell runs the full 60-case suite, producing trajectory traces and per-cell metrics. The run_full.py orchestrator handles checkpointing, retries, and parallel execution across --workers.

Rubric-Scoring Layer: Deterministic Multi-Dimensional Grading

The rubric—defined in book/chapter7.md (lines 25–58)—enforces four graded dimensions plus a safety veto:

Dimension Weight Description
Precision essential Correctness of recalled facts
Recall important Coverage of relevant memories
Reasoning important Quality of explanation for selections
Proactivity important Appropriate initiative in ambiguous queries
Hallucination Veto veto Any fabricated content zeroes the score

Weights are explicitly marked (essential, important, veto), making the scoring auditable and resistant to gaming.


Automated Backend Substitution for Benchmark Stability

External API providers fail or rotate credentials. Rather than crash, the framework substitutes equivalent backends automatically.

Substitution rules live in default_config.yaml. When SiliconFlow is unavailable, the matrix transparently routes to OpenRouter's baai/bge-m3 equivalent. Every substitution is logged to results/full_matrix_backend_readiness_20260731.json (see README lines 42–58), preserving experimental provenance.

This ensures benchmarks remain runnable without manual intervention, a critical property for longitudinal studies or CI/CD pipelines.


Running the Full Evaluation Matrix

Experiment 7-11: Complete 60-Case Run

cd chapter7/user-memory-system-evaluation
python -m pip install -r requirements.txt
cp env.example .env      # add required API keys

python run_full.py 7-11 \
    --config default_config.yaml \
    --workers 4 \
    --readiness results/full_matrix_backend_readiness.json \
    --output results/full_7_11_60_case_matrix.json

The orchestrator probes all backends, substitutes where needed, executes 60 cases per matrix cell, and writes a unified JSON report with hit@5, recall@5, MRR, step counts, tool-call statistics, and cost breakdowns.

Single Configuration Testing (Experiment 7-4)

For rapid iteration on one backend combination:

python experiment.py 7-4 \
    --config default_config.yaml \
    --output results/experiment_7_4.json

Loading and Inspecting Rubric Data

Programmatic Rubric Access

The rubric YAML block can be extracted from the book source for custom judges:

import yaml, json, pathlib

rubric_path = pathlib.Path("../book/chapter7.md")
with rubric_path.open() as f:
    for line in f:
        if line.strip().startswith("rubric:"):
            yaml_block = "\n".join(f.readlines()[0:30])
            rubric = yaml.safe_load(yaml_block)
            break

print(json.dumps(rubric, indent=2, ensure_ascii=False))

Reading Per-Cell Results

import json, pathlib

data = json.loads(
    pathlib.Path("results/full_7_11_60_case_matrix.json").read_text()
)

# Extract Kimi + BGE-M3 + no reranker score

score = data["Kimi"]["bge-m3"]["no-reranker"]["rubric_score"]
print("Rubric score:", score)

The evidence file results/full_7_3_structured_rubric_evidence.json contains the complete four-grade judgments with hallucination vetos for all 60 cases, enabling independent verification via validation/verify_full_matrix_20260731.py.


Design Principles for Extensible Benchmarks

The ai-agent-book framework embodies four patterns transferable to any agent evaluation domain:

  1. Curated, frozen test cases — prevents test-set contamination and enables fair comparison
  2. Explicit matrix enumeration — makes the evaluation surface complete and inspectable
  3. Structured rubrics with veto powers — encodes domain knowledge and safety constraints directly
  4. Backend abstraction with substitution logging — decouples results from provider availability

These principles are discussed in the multi-modal LLM-as-a-Judge section of chapter 7, suggesting extension to vision, audio, or tool-use agents.


Summary

  • Three-layer architecture: test cases, system matrix, and rubric scoring create reproducible, auditable benchmarks
  • 60 curated user-memory scenarios provide a stable evaluation surface across all experiments
  • Four-dimensional rubric (precision, recall, reasoning, proactivity) plus hallucination veto enforces deterministic, safety-conscious grading
  • Automatic backend substitution with logged decisions keeps benchmarks runnable despite provider instability
  • Complete artifact chain: from default_config.yaml through run_full.py to full_7_11_60_case_matrix.json and independent verify_full_matrix_20260731.py validation

Frequently Asked Questions

How does the hallucination veto work in practice?

Any detected fabrication automatically zeros the rubric score regardless of other dimensions. This is implemented in the Experiment 7-3 judge according to lines 25–58 of book/chapter7.md. The veto appears as a boolean field in results/full_7_3_structured_rubric_evidence.json, making hallucination incidents directly auditable.

Can I add new embeddings or models without rewriting test cases?

Yes. The matrix in default_config.yaml is designed for extension. Add new entries to the embeddings, rerankers, or models lists, and run_full.py will automatically include them in the full factorial run. The 60 test cases in chapter3/user-memory-evaluation/test_cases remain unchanged, ensuring comparability.

What happens if an API provider is temporarily unavailable?

The framework probes all configured backends and substitutes equivalent alternatives per rules in default_config.yaml. Substitutions are logged to results/full_matrix_backend_readiness_20260731.json with timestamps and rationale, preserving experimental reproducibility even when original providers fail.

How do I verify that the rubric was applied consistently?

Run validation/verify_full_matrix_20260731.py against results/full_7_3_structured_rubric_evidence.json. This independent checker re-evaluates a sample of judgments against the rubric definition to detect scoring drift or implementation bugs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →