How to Evaluate Agent Performance: Metrics, Benchmarks, and Significance in the AI-Agent-Book Repository

The AI-Agent-Book repository implements a three-tier evaluation system using behavior-level metrics, update-level metrics, and dual-benchmark comparison to determine whether an agent modification is safe to deploy.

This guide explains the complete evaluation framework used in the bojieli/ai-agent-book project to quantify autonomous agent performance. The system combines sandboxed execution, operational quality checks, and statistical significance testing against stable baselines.

Behavior-Level Metrics: Quantifying Raw Agent Behavior

The behavior_metrics function in [chapter8/self-modifying-agent/evolution.py](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-modifying-agent/evolution.py) (lines 46–66) executes candidates inside a locked-down sandbox and returns four critical measurements:

Metric Description Target Value
mean_nonretryable_calls Average fatal API errors per run Lower is better (ideally 0)
temporary_error_recovery_rate Success rate at recovering from transient failures Higher is better (ideally 1.0)
old_task_regressions Number of previously-solved tasks now failing Zero
evaluation_failed Boolean flag for sandbox crashes False

These metrics answer a fundamental question: Does the candidate agent crash less and recover faster than before?

Measuring Behavior Metrics in Practice

from chapter8.self_modifying_agent.evolution import behavior_metrics

# trajectories: list of execution logs from previous runs

# candidate_source: string containing the new agent source code

metrics = behavior_metrics(candidate_source, trajectories)

print(metrics)

# {

#   "mean_nonretryable_calls": 1.0,

#   "temporary_error_recovery_rate": 1.0,

#   "old_task_regressions": 0,

#   "evaluation_failed": False,

# }

The sandbox runner isolates each candidate execution to prevent side effects from corrupting the evaluation environment.

Update-Level Metrics: Validating Modification Quality

When an agent proposes self-modifications, raw behavioral stability is insufficient. The [harness.py](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-evolution-eval/harness.py) module (lines 124–128) collects three operational metrics that verify the modification itself is sound:

  • candidate_modification_validity — Does the patch respect syntactic constraints and authorized imports?
  • artifact_activation_rate — Does the new code actually execute during agent runs?
  • memory_adherence_rate — Does the agent respect memory constraints after the change?

These metrics distinguish between behavior that looks good and modifications that work correctly in production.

Running Full Evaluation with Update Metrics

from chapter8.self_evolution_eval.harness import evaluate_candidate
from pathlib import Path

result = evaluate_candidate(
    candidate_path=Path("candidate.py"),
    stable_path=Path("stable.py"),
    agent=agent,  # instantiated Agent object

)

behavior = result["metrics"]         # sandbox behavior metrics

updates = result["update_metrics"]   # activation & adherence rates

Benchmark Comparison: Establishing Statistical Significance

The repository uses two reference benchmarks to judge candidate quality:

  1. Stable baseline — The previous trusted version of the agent
  2. Real-LLM reference — A gold-standard implementation expected to perform optimally

A candidate achieves significance only when it satisfies all four conditions simultaneously:

  • Reduces mean_nonretryable_calls relative to baseline
  • Matches or exceeds real-LLM temporary_error_recovery_rate
  • Registers zero old_task_regressions
  • Achieves non-null, sufficiently high update-level metrics

This logic appears in experimental scripts like [run_experiment_8_5.py](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/self-modifying-agent/run_experiment_8_5.py):


# Simplified benchmark comparison logic

def is_significant(candidate_metrics, baseline_metrics, real_llm_metrics):
    return (
        candidate_metrics["mean_nonretryable_calls"] < baseline_metrics["mean_nonretryable_calls"]
        and candidate_metrics["temporary_error_recovery_rate"] >= real_llm_metrics["temporary_error_recovery_rate"]
        and candidate_metrics["old_task_regressions"] == 0
    )

The release_manifest logic in evolution.py (lines 76–84) enforces these thresholds — candidates failing any condition are rejected.

Complete Evaluation Pipeline

The full workflow for evaluating agent performance follows five stages:

  1. Generate trajectories — Execute stable baseline, real-LLM, and candidate to produce execution logs
  2. Compute behavior metrics — Run behavior_metrics() in sandbox for each version
  3. Collect update metrics — Execute candidate through harness to measure activation and adherence
  4. Benchmark comparison — Apply significance logic against both reference points
  5. Generate release manifest — Record final decision (release_to_canary or reject_candidate)

This architecture ensures reproducible, automated, and safe agent evolution.

Summary

  • Behavior-level metrics in evolution.py quantify sandbox performance: error rates, recovery rates, and regression counts
  • Update-level metrics in harness.py validate modification quality: validity, activation, and memory adherence
  • Dual-benchmark comparison requires candidates to improve on baseline while matching real-LLM performance
  • Significance criteria demand zero regressions, reduced fatal errors, sustained recovery rates, and valid modifications
  • All evaluations execute in isolated sandboxes with automated release decisions

Frequently Asked Questions

How does the sandbox prevent evaluation corruption?

The sandbox_runner.py module creates isolated execution environments where each candidate runs with restricted permissions. If a candidate crashes the sandbox, evaluation_failed becomes True and other metrics return None, preventing corrupted data from influencing release decisions.

Why compare against both baseline and real-LLM benchmarks?

The stable baseline establishes improvement over the current production version, while the real-LLM reference sets an absolute performance ceiling. A candidate must satisfy both: improving on deployed code without falling below theoretical best-case performance. This prevents accepting regressions masked by coincidental baseline weaknesses.

What happens when old_task_regressions is non-zero?

Any positive value triggers automatic rejection. The release_manifest logic in evolution.py treats task regression as a hard failure condition — even if all other metrics improve, forgetting previously mastered capabilities violates the safety constraints of the self-modifying system.

Can update-level metrics be evaluated without the full harness?

No. While behavior metrics require only the sandbox, update metrics demand integration with the complete agent system to verify that modifications actually activate and respect operational constraints. The harness provides this full-system context that isolated sandbox execution cannot replicate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →