# How to Evaluate Agent Performance: Metrics, Benchmarks, and Significance in the AI-Agent-Book Repository

> Discover how to evaluate agent performance with key metrics, benchmarks, and significance. Learn the three-tier system for safe AI agent deployment from the AI-Agent-Book repository.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-06

---

**The AI-Agent-Book repository implements a three-tier evaluation system using behavior-level metrics, update-level metrics, and dual-benchmark comparison to determine whether an agent modification is safe to deploy.**

This guide explains the complete evaluation framework used in the `bojieli/ai-agent-book` project to quantify autonomous agent performance. The system combines **sandboxed execution**, **operational quality checks**, and **statistical significance testing** against stable baselines.

## Behavior-Level Metrics: Quantifying Raw Agent Behavior

The `behavior_metrics` function in [[`chapter8/self-modifying-agent/evolution.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-modifying-agent/evolution.py)](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-modifying-agent/evolution.py) (lines 46–66) executes candidates inside a **locked-down sandbox** and returns four critical measurements:

| Metric | Description | Target Value |
|--------|-------------|------------|
| `mean_nonretryable_calls` | Average fatal API errors per run | **Lower is better** (ideally 0) |
| `temporary_error_recovery_rate` | Success rate at recovering from transient failures | **Higher is better** (ideally 1.0) |
| `old_task_regressions` | Number of previously-solved tasks now failing | **Zero** |
| `evaluation_failed` | Boolean flag for sandbox crashes | **False** |

These metrics answer a fundamental question: *Does the candidate agent crash less and recover faster than before?*

### Measuring Behavior Metrics in Practice

```python
from chapter8.self_modifying_agent.evolution import behavior_metrics

# trajectories: list of execution logs from previous runs

# candidate_source: string containing the new agent source code

metrics = behavior_metrics(candidate_source, trajectories)

print(metrics)

# {

#   "mean_nonretryable_calls": 1.0,

#   "temporary_error_recovery_rate": 1.0,

#   "old_task_regressions": 0,

#   "evaluation_failed": False,

# }

```

The sandbox runner isolates each candidate execution to prevent side effects from corrupting the evaluation environment.

## Update-Level Metrics: Validating Modification Quality

When an agent proposes self-modifications, raw behavioral stability is insufficient. The [[`harness.py`](https://github.com/bojieli/ai-agent-book/blob/main/harness.py)](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-evolution-eval/harness.py) module (lines 124–128) collects **three operational metrics** that verify the modification itself is sound:

- **`candidate_modification_validity`** — Does the patch respect syntactic constraints and authorized imports?
- **`artifact_activation_rate`** — Does the new code actually execute during agent runs?
- **`memory_adherence_rate`** — Does the agent respect memory constraints after the change?

These metrics distinguish between *behavior that looks good* and *modifications that work correctly in production*.

### Running Full Evaluation with Update Metrics

```python
from chapter8.self_evolution_eval.harness import evaluate_candidate
from pathlib import Path

result = evaluate_candidate(
    candidate_path=Path("candidate.py"),
    stable_path=Path("stable.py"),
    agent=agent,  # instantiated Agent object

)

behavior = result["metrics"]         # sandbox behavior metrics

updates = result["update_metrics"]   # activation & adherence rates

```

## Benchmark Comparison: Establishing Statistical Significance

The repository uses **two reference benchmarks** to judge candidate quality:

1. **Stable baseline** — The previous trusted version of the agent
2. **Real-LLM reference** — A gold-standard implementation expected to perform optimally

A candidate achieves **significance** only when it satisfies all four conditions simultaneously:

- Reduces `mean_nonretryable_calls` relative to baseline
- Matches or exceeds real-LLM `temporary_error_recovery_rate`
- Registers **zero** `old_task_regressions`
- Achieves non-null, sufficiently high update-level metrics

This logic appears in experimental scripts like [[`run_experiment_8_5.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_experiment_8_5.py)](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-modifying-agent/run_experiment_8_5.py):

```python

# Simplified benchmark comparison logic

def is_significant(candidate_metrics, baseline_metrics, real_llm_metrics):
    return (
        candidate_metrics["mean_nonretryable_calls"] < baseline_metrics["mean_nonretryable_calls"]
        and candidate_metrics["temporary_error_recovery_rate"] >= real_llm_metrics["temporary_error_recovery_rate"]
        and candidate_metrics["old_task_regressions"] == 0
    )

```

The `release_manifest` logic in [`evolution.py`](https://github.com/bojieli/ai-agent-book/blob/main/evolution.py) (lines 76–84) enforces these thresholds — candidates failing any condition are rejected.

## Complete Evaluation Pipeline

The full workflow for evaluating agent performance follows five stages:

1. **Generate trajectories** — Execute stable baseline, real-LLM, and candidate to produce execution logs
2. **Compute behavior metrics** — Run `behavior_metrics()` in sandbox for each version
3. **Collect update metrics** — Execute candidate through harness to measure activation and adherence
4. **Benchmark comparison** — Apply significance logic against both reference points
5. **Generate release manifest** — Record final decision (`release_to_canary` or `reject_candidate`)

This architecture ensures **reproducible, automated, and safe** agent evolution.

## Summary

- **Behavior-level metrics** in [`evolution.py`](https://github.com/bojieli/ai-agent-book/blob/main/evolution.py) quantify sandbox performance: error rates, recovery rates, and regression counts
- **Update-level metrics** in [`harness.py`](https://github.com/bojieli/ai-agent-book/blob/main/harness.py) validate modification quality: validity, activation, and memory adherence
- **Dual-benchmark comparison** requires candidates to improve on baseline while matching real-LLM performance
- **Significance criteria** demand zero regressions, reduced fatal errors, sustained recovery rates, and valid modifications
- All evaluations execute in isolated sandboxes with automated release decisions

## Frequently Asked Questions

### How does the sandbox prevent evaluation corruption?

The [`sandbox_runner.py`](https://github.com/bojieli/ai-agent-book/blob/main/sandbox_runner.py) module creates isolated execution environments where each candidate runs with restricted permissions. If a candidate crashes the sandbox, `evaluation_failed` becomes `True` and other metrics return `None`, preventing corrupted data from influencing release decisions.

### Why compare against both baseline and real-LLM benchmarks?

The **stable baseline** establishes improvement over the current production version, while the **real-LLM reference** sets an absolute performance ceiling. A candidate must satisfy both: improving on deployed code without falling below theoretical best-case performance. This prevents accepting regressions masked by coincidental baseline weaknesses.

### What happens when old_task_regressions is non-zero?

Any positive value triggers automatic rejection. The `release_manifest` logic in [`evolution.py`](https://github.com/bojieli/ai-agent-book/blob/main/evolution.py) treats task regression as a hard failure condition — even if all other metrics improve, forgetting previously mastered capabilities violates the safety constraints of the self-modifying system.

### Can update-level metrics be evaluated without the full harness?

No. While behavior metrics require only the sandbox, update metrics demand integration with the complete agent system to verify that modifications actually activate and respect operational constraints. The harness provides this full-system context that isolated sandbox execution cannot replicate.