# Key AI Agent Evaluation Environments in the ai-agent-book Repository

> Discover key AI agent evaluation environments in the ai-agent-book repo. Test autonomous agents in health data analysis, world simulation, sandboxing, and browser automation with reproducible scoring.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: article
- Published: 2026-08-24

---

**The ai-agent-book repository implements five production-ready AI agent evaluation environments—ReportingEnvironment, GAIA Experience Environment, AWorld Sandbox, Trajectory-Verifier Layer, and Browser-Use RPA Environment—that provide deterministic APIs and reproducible scoring for testing autonomous agents across public health data analysis, multimodal world simulation, OS-level sandboxing, and browser automation tasks.**

The `bojieli/ai-agent-book` repository provides a comprehensive collection of **AI agent evaluation environments** designed to simulate real-world contexts for rigorous autonomous system testing. Each environment supplies deterministic data sources, well-defined tool APIs, and structured evidence tracking to ensure reproducible evaluation of agent reasoning and execution capabilities.

## ReportingEnvironment: Public Health Data Simulation

The **ReportingEnvironment** serves as a deterministic testbed for public health reporting scenarios, implemented in [`chapter7/public-health-reporting-eval/reporting_tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter7/public-health-reporting-eval/reporting_tools.py). This environment provides a synthetic DHIS2-style health dataset loaded via a CSV loader that normalizes integer fields, enabling agents to perform auditable data quality assessments.

Agents interact with the environment through a `_select` query helper and specialized tool methods including `calculate_test_positivity`, `calculate_reporting_completeness`, `compare_confirmed_cases`, `find_data_quality_issues`, and `review_stockouts`. When an agent invokes `environment.call(tool, args)`, the environment returns structured health metrics alongside an **evidence** field that points to the specific raw rows justifying the calculation, ensuring full reproducibility of reasoning chains.

```python
from reporting_tools import ReportingEnvironment

# Initialise with the synthetic CSV file

env = ReportingEnvironment("data/synthetic_reports.csv")

# Ask the environment to run a tool

result = env.call(
    "calculate_test_positivity",
    {"org_unit_id": "OU_001", "period": "2023Q1"}
)

print(result)

# ► {'tests': 1200, 'confirmed_cases': 30, 'test_positivity_pct': 2.5,

#    'evidence': ['row_001']}

```

## GAIA Experience Environment: Multimodal World Simulation

According to the `bojieli/ai-agent-book` source code, the **GAIA Experience Environment** simulates a rich, multimodal world where agents interact with facilities and receive observations scored on task completion. The environment configuration is defined in `env.template`, which specifies the JSON schema for the world state, while [`experience_agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/experience_agent.py) and [`experience_documents.py`](https://github.com/bojieli/ai-agent-book/blob/main/experience_documents.py) generate and manage agent trajectories.

The [`llm_env.py`](https://github.com/bojieli/ai-agent-book/blob/main/llm_env.py) wrapper exposes this environment to LLM-driven agents, producing **environment scores** that serve as evidence for downstream evaluation layers such as the trajectory-verifier. This architecture enables end-to-end assessment of perception, planning, and execution capabilities in a controlled simulated world.

```python
from gaia_experience.experience_agent import GAIAAgent
from gaia_experience.llm_env import GAIAEnvironment

# Load a predefined world configuration

env = GAIAEnvironment("config.yaml")
agent = GAIAAgent(env)

# Run a single episode

trajectory = agent.run_episode(task="answer_user_query")
print(trajectory.final_score)   # numeric environment score

```

## AWorld Sandbox and OSWorld Multi-Env: Containerized OS Simulation

For agents requiring full operating system access, the **AWorld Sandbox** and **OSWorld Multi-Env** provide isolated Docker-based simulations. As implemented in `chapter9/gaia-experience/AWorld`, this environment layer is orchestrated by [`run_multienv_aworldAgent.py`](https://github.com/bojieli/ai-agent-book/blob/main/run_multienv_aworldAgent.py), which can spin up multiple sandbox instances simultaneously for parallel evaluation.

The abstract sandbox API defined in [`base_sandbox_api.py`](https://github.com/bojieli/ai-agent-book/blob/main/base_sandbox_api.py) exposes three critical lifecycle methods: `reset`, `step`, and `close`, while [`common.py`](https://github.com/bojieli/ai-agent-book/blob/main/common.py) handles environment-variable substitution and logging. This sandbox architecture delivers a **realistic runtime** for tasks such as phone-use and browser-use experiments while maintaining determinism through snapshot and restore functionality.

```python
from AWorld.examples.osworld.run_multienv_aworldAgent import launch_agents

# Launch three sandbox instances, each receiving its own config

launch_agents(num_envs=3, config_base_dir="evaluation_examples")

```

## Trajectory-Verifier Environment Layer: State Validation

The **Trajectory-Verifier Environment Layer** functions as a lightweight policy-and-environment checker that validates final trajectory states against expected outcomes. Located in [`chapter9/trajectory-verifier/verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/trajectory-verifier/verifier.py), this layer defines the `environment_result` evaluation dimension that examines whether tools succeeded and agents achieved their objectives.

The verifier maps discrete outcomes—**PASS**, **FAIL**, and **UNCERTAIN**—to numeric scores, creating a calibrated evaluation metric. This layer integrates with complementary evaluation dimensions (policy, rubric) to form a multi-dimensional assessment matrix that quantifies task completion accuracy.

```python
from chapter9.trajectory_verifier.verifier import Verifier

verifier = Verifier()

# `report` is a list of trajectory dicts produced by an agent

scores = verifier.evaluate(report, dimensions=["environment_result"])
print(scores)

# ► {'environment_result': {'PASS': 0.85, 'FAIL': 0.15}}

```

## Browser-Use RPA Environment: Headless Browser Automation

The **Browser-Use RPA Environment** emulates deterministic headless browser sessions for testing web navigation and form-filling capabilities. The CLI entry point in [`chapter9/browser-use-rpa/browser-use/browser_use/cli.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/browser-use-rpa/browser-use/browser_use/cli.py) orchestrates runs and logs `evaluation_previous_goal` entries, while [`events.py`](https://github.com/bojieli/ai-agent-book/blob/main/events.py) defines environment-specific events including timeouts and navigation success states.

Agent outputs conform to the JSON schema defined in [`browser_use/agent/views.py`](https://github.com/bojieli/ai-agent-book/blob/main/browser_use/agent/views.py), which includes the **`evaluation_previous_goal`** field capturing the environment's assessment of whether the previous step was achieved. This design allows precise measurement of incremental progress in multi-step web automation tasks.

```bash

# From the repository root

python -m chapter9.browser-use-rpa.browser-use.browser_use.cli \
    --task "fill_health_form" \
    --output result.json

```

## Shared Architectural Patterns

All **AI agent evaluation environments** in the repository follow a consistent four-pattern design that ensures interoperability and reproducibility:

- **Deterministic data sources** (CSV files, JSON configurations, or Docker snapshots) eliminate variability between test runs.
- **Explicit APIs** (`call`, `step`, `reset`, `close`) provide standardized interfaces for agent interaction across all environment types.
- **Evidence tracking** requires each tool to return an `evidence` field pointing to raw data rows or snapshots that justify the answer, enabling audit trails.
- **Scoring hooks** allow higher-level evaluation frameworks to consume evidence and compute task-level metrics including accuracy, completeness, and safety.

This standardized architecture enables researchers to swap environments without modifying agent implementations, facilitating cross-domain stress testing of autonomous systems.

## Summary

- The **ReportingEnvironment** provides deterministic public health data simulation with built-in data quality tools for testing analytical agent capabilities.
- The **GAIA Experience Environment** offers multimodal world simulation with trajectory scoring for end-to-end perception and planning evaluation.
- The **AWorld Sandbox** delivers containerized OS-level environments via Docker, supporting realistic phone-use and browser-use scenarios with snapshot-based determinism.
- The **Trajectory-Verifier Layer** implements state validation and PASS/FAIL/UNCERTAIN scoring to quantify task completion accuracy.
- All environments share a unified design pattern emphasizing deterministic data sources, explicit APIs, evidence tracking, and pluggable scoring hooks.

## Frequently Asked Questions

### What makes the ai-agent-book evaluation environments deterministic?

Each environment uses fixed initial states—whether CSV datasets in `ReportingEnvironment`, JSON templates in GAIA Experience, or Docker snapshots in AWorld Sandbox—ensuring that identical agent actions produce identical outcomes across repeated evaluations. This determinism is critical for reproducible AI research and valid performance comparisons.

### How does the Trajectory-Verifier Layer score agent performance?

The layer defined in [`verifier.py`](https://github.com/bojieli/ai-agent-book/blob/main/verifier.py) examines the final state of agent trajectories and assigns discrete labels (**PASS**, **FAIL**, or **UNCERTAIN**) based on whether the agent successfully invoked tools and achieved objectives. These labels are mapped to numeric scores to provide calibrated performance metrics that integrate with multi-dimensional evaluation matrices.

### What is the difference between the GAIA Experience Environment and the AWorld Sandbox?

The **GAIA Experience Environment** focuses on high-level multimodal simulation with structured world states and task-completion scoring, while the **AWorld Sandbox** provides low-level OS access through Docker containers for scenarios requiring actual file system, network, or application interactions. GAIA is used for abstract reasoning evaluation, whereas AWorld supports concrete tool-use scenarios like phone automation.

### How does evidence tracking work in these evaluation environments?

When an agent invokes a tool via `environment.call()` or similar APIs, the environment returns not only the result but also an `evidence` field containing references to specific data rows, snapshots, or state identifiers that justified the answer. This mechanism enables researchers to audit agent decisions and verify that conclusions are grounded in actual environment states rather than hallucinations.