Key AI Agent Evaluation Environments in the ai-agent-book Repository
The ai-agent-book repository implements five production-ready AI agent evaluation environments—ReportingEnvironment, GAIA Experience Environment, AWorld Sandbox, Trajectory-Verifier Layer, and Browser-Use RPA Environment—that provide deterministic APIs and reproducible scoring for testing autonomous agents across public health data analysis, multimodal world simulation, OS-level sandboxing, and browser automation tasks.
The bojieli/ai-agent-book repository provides a comprehensive collection of AI agent evaluation environments designed to simulate real-world contexts for rigorous autonomous system testing. Each environment supplies deterministic data sources, well-defined tool APIs, and structured evidence tracking to ensure reproducible evaluation of agent reasoning and execution capabilities.
ReportingEnvironment: Public Health Data Simulation
The ReportingEnvironment serves as a deterministic testbed for public health reporting scenarios, implemented in chapter7/public-health-reporting-eval/reporting_tools.py. This environment provides a synthetic DHIS2-style health dataset loaded via a CSV loader that normalizes integer fields, enabling agents to perform auditable data quality assessments.
Agents interact with the environment through a _select query helper and specialized tool methods including calculate_test_positivity, calculate_reporting_completeness, compare_confirmed_cases, find_data_quality_issues, and review_stockouts. When an agent invokes environment.call(tool, args), the environment returns structured health metrics alongside an evidence field that points to the specific raw rows justifying the calculation, ensuring full reproducibility of reasoning chains.
from reporting_tools import ReportingEnvironment
# Initialise with the synthetic CSV file
env = ReportingEnvironment("data/synthetic_reports.csv")
# Ask the environment to run a tool
result = env.call(
"calculate_test_positivity",
{"org_unit_id": "OU_001", "period": "2023Q1"}
)
print(result)
# ► {'tests': 1200, 'confirmed_cases': 30, 'test_positivity_pct': 2.5,
# 'evidence': ['row_001']}
GAIA Experience Environment: Multimodal World Simulation
According to the bojieli/ai-agent-book source code, the GAIA Experience Environment simulates a rich, multimodal world where agents interact with facilities and receive observations scored on task completion. The environment configuration is defined in env.template, which specifies the JSON schema for the world state, while experience_agent.py and experience_documents.py generate and manage agent trajectories.
The llm_env.py wrapper exposes this environment to LLM-driven agents, producing environment scores that serve as evidence for downstream evaluation layers such as the trajectory-verifier. This architecture enables end-to-end assessment of perception, planning, and execution capabilities in a controlled simulated world.
from gaia_experience.experience_agent import GAIAAgent
from gaia_experience.llm_env import GAIAEnvironment
# Load a predefined world configuration
env = GAIAEnvironment("config.yaml")
agent = GAIAAgent(env)
# Run a single episode
trajectory = agent.run_episode(task="answer_user_query")
print(trajectory.final_score) # numeric environment score
AWorld Sandbox and OSWorld Multi-Env: Containerized OS Simulation
For agents requiring full operating system access, the AWorld Sandbox and OSWorld Multi-Env provide isolated Docker-based simulations. As implemented in chapter9/gaia-experience/AWorld, this environment layer is orchestrated by run_multienv_aworldAgent.py, which can spin up multiple sandbox instances simultaneously for parallel evaluation.
The abstract sandbox API defined in base_sandbox_api.py exposes three critical lifecycle methods: reset, step, and close, while common.py handles environment-variable substitution and logging. This sandbox architecture delivers a realistic runtime for tasks such as phone-use and browser-use experiments while maintaining determinism through snapshot and restore functionality.
from AWorld.examples.osworld.run_multienv_aworldAgent import launch_agents
# Launch three sandbox instances, each receiving its own config
launch_agents(num_envs=3, config_base_dir="evaluation_examples")
Trajectory-Verifier Environment Layer: State Validation
The Trajectory-Verifier Environment Layer functions as a lightweight policy-and-environment checker that validates final trajectory states against expected outcomes. Located in chapter9/trajectory-verifier/verifier.py, this layer defines the environment_result evaluation dimension that examines whether tools succeeded and agents achieved their objectives.
The verifier maps discrete outcomes—PASS, FAIL, and UNCERTAIN—to numeric scores, creating a calibrated evaluation metric. This layer integrates with complementary evaluation dimensions (policy, rubric) to form a multi-dimensional assessment matrix that quantifies task completion accuracy.
from chapter9.trajectory_verifier.verifier import Verifier
verifier = Verifier()
# `report` is a list of trajectory dicts produced by an agent
scores = verifier.evaluate(report, dimensions=["environment_result"])
print(scores)
# ► {'environment_result': {'PASS': 0.85, 'FAIL': 0.15}}
Browser-Use RPA Environment: Headless Browser Automation
The Browser-Use RPA Environment emulates deterministic headless browser sessions for testing web navigation and form-filling capabilities. The CLI entry point in chapter9/browser-use-rpa/browser-use/browser_use/cli.py orchestrates runs and logs evaluation_previous_goal entries, while events.py defines environment-specific events including timeouts and navigation success states.
Agent outputs conform to the JSON schema defined in browser_use/agent/views.py, which includes the evaluation_previous_goal field capturing the environment's assessment of whether the previous step was achieved. This design allows precise measurement of incremental progress in multi-step web automation tasks.
# From the repository root
python -m chapter9.browser-use-rpa.browser-use.browser_use.cli \
--task "fill_health_form" \
--output result.json
Shared Architectural Patterns
All AI agent evaluation environments in the repository follow a consistent four-pattern design that ensures interoperability and reproducibility:
- Deterministic data sources (CSV files, JSON configurations, or Docker snapshots) eliminate variability between test runs.
- Explicit APIs (
call,step,reset,close) provide standardized interfaces for agent interaction across all environment types. - Evidence tracking requires each tool to return an
evidencefield pointing to raw data rows or snapshots that justify the answer, enabling audit trails. - Scoring hooks allow higher-level evaluation frameworks to consume evidence and compute task-level metrics including accuracy, completeness, and safety.
This standardized architecture enables researchers to swap environments without modifying agent implementations, facilitating cross-domain stress testing of autonomous systems.
Summary
- The ReportingEnvironment provides deterministic public health data simulation with built-in data quality tools for testing analytical agent capabilities.
- The GAIA Experience Environment offers multimodal world simulation with trajectory scoring for end-to-end perception and planning evaluation.
- The AWorld Sandbox delivers containerized OS-level environments via Docker, supporting realistic phone-use and browser-use scenarios with snapshot-based determinism.
- The Trajectory-Verifier Layer implements state validation and PASS/FAIL/UNCERTAIN scoring to quantify task completion accuracy.
- All environments share a unified design pattern emphasizing deterministic data sources, explicit APIs, evidence tracking, and pluggable scoring hooks.
Frequently Asked Questions
What makes the ai-agent-book evaluation environments deterministic?
Each environment uses fixed initial states—whether CSV datasets in ReportingEnvironment, JSON templates in GAIA Experience, or Docker snapshots in AWorld Sandbox—ensuring that identical agent actions produce identical outcomes across repeated evaluations. This determinism is critical for reproducible AI research and valid performance comparisons.
How does the Trajectory-Verifier Layer score agent performance?
The layer defined in verifier.py examines the final state of agent trajectories and assigns discrete labels (PASS, FAIL, or UNCERTAIN) based on whether the agent successfully invoked tools and achieved objectives. These labels are mapped to numeric scores to provide calibrated performance metrics that integrate with multi-dimensional evaluation matrices.
What is the difference between the GAIA Experience Environment and the AWorld Sandbox?
The GAIA Experience Environment focuses on high-level multimodal simulation with structured world states and task-completion scoring, while the AWorld Sandbox provides low-level OS access through Docker containers for scenarios requiring actual file system, network, or application interactions. GAIA is used for abstract reasoning evaluation, whereas AWorld supports concrete tool-use scenarios like phone automation.
How does evidence tracking work in these evaluation environments?
When an agent invokes a tool via environment.call() or similar APIs, the environment returns not only the result but also an evidence field containing references to specific data rows, snapshots, or state identifiers that justified the answer. This mechanism enables researchers to audit agent decisions and verify that conclusions are grounded in actual environment states rather than hallucinations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →