How Experiments Are Organized in the AI Agent Book: A Technical Guide

Experiments in the AI Agent Book follow a strict modular structure with three distinct scripts—experiment.py for exploration, run_experiment_X_Y.py for canonical evidence generation, and finalize_experiment_X_Y.py for post-processing—organized under chapter-specific directories with versioned validation artifacts.

The bojieli/ai-agent-book repository implements a rigorous, reproducible research framework where every experiment is encapsulated as a self-contained module. Understanding how experiments are organized in the AI Agent Book reveals a deliberate architecture that separates exploratory iteration from acceptance-grade evidence collection, ensuring every figure and table in the text can be traced to serialized raw data.

The Three-Tier Experiment Architecture

Each experiment directory contains a standardized trio of Python scripts that serve distinct phases of the research workflow. This separation prevents exploratory noise from contaminating citable results while allowing rapid iteration during development.

Exploratory Runner (experiment.py)

The experiment.py file serves as the flexible CLI entry point for iterative development. Located at paths like chapter1/learning-from-experience/experiment.py, this script supports dynamic parameter switching between Q-learning and LLM agents, custom episode counts, and interactive plotting capabilities.

Researchers use this script to explore hyperparameters, generate debugging plots, or compare model behaviors without strict protocol enforcement. The CLI accepts flags such as --mode (accepting values like both, rl-only, or llm-only), --rl-episodes, --llm-episodes, and --model (e.g., kimi-k3).

Canonical Runner (run_experiment_X_Y.py)

For deterministic, book-citable results, the repository provides canonical runners such as run_experiment_7_2.py and run_experiment_9_1.py. These scripts enforce strict protocols: fixed episode counts (e.g., 10,000 RL episodes and 20 LLM episodes for Experiment 7-2), mandatory raw response logging, and abort-on-error handling to ensure every run produces valid evidence.

The canonical runner calls underlying functions like run_full_comparison() with hardcoded parameters and writes timestamped output to validation/<timestamp>/evidence.json. Any deviation—such as missing response IDs, fallback parsers triggering, or API errors—causes immediate termination to preserve evidence integrity.

Evidence Finalizer (finalize_experiment_X_Y.py)

When serialization fails after successful LLM API calls, finalize_experiment_7_2.py regenerates the evidence JSON without incurring additional API costs. This script loads existing raw campaign data from the validation directory, validates it against the required schema, and writes a clean evidence.json that the book references directly.

Directory Structure and File Organization

Experiments are nested within chapter directories following a predictable schema. Each chapter lives under chapter<N>/ and contains topic-specific experiment directories (e.g., learning-from-experience, context-compression, trajectory-verifier).

The typical layout for an experiment directory includes:

chapter1/learning-from-experience/
├── experiment.py                # Flexible CLI runner

├── run_experiment_7_2.py       # Canonical runner (Experiment 7-2)

├── finalize_experiment_7_2.py  # Evidence-only finalizer

├── env.example                 # API key template

├── game_environment.py         # Environment logic

├── rl_agent.py                 # Q-learning implementation

├── llm_agent.py                # LLM agent implementation

├── tests/                      # Unit and integration tests

│   ├── test_basic.py
│   └── ...
├── validation/                 # Serialized evidence

│   └── 20260730_011704/
│       └── evidence.json
└── README.md                   # Human overview (lists experiments 7-1, 7-2)

The README.md in each experiment directory provides human-readable context. For example, the Chapter 1 README explicitly lists Experiment 7-1 (Q-learning only) and Experiment 7-2 (full RL vs LLM comparison), along with quick-start commands for replication.

The Experiment Lifecycle

The organization of files supports a four-phase workflow that isolates exploration from publication-grade evidence generation:

  1. Exploratory Phase – Developers run experiment.py with custom flags (e.g., --mode both --model kimi-k3) to explore hyperparameters, generate plots, or debug agents without protocol constraints.

  2. Canonical Phase – When the experimental protocol is finalized, running run_experiment_X_Y.py executes the exact episode counts and logging requirements specified in the book text. This generates the official evidence used for citations.

  3. Evidence Generation – The canonical script writes a timestamped evidence.json under validation/. This file contains serialized raw LLM responses, timestamps, and metrics that serve as the formal proof for the experiment.

  4. Post-Processing – If only the serialization step fails after a successful LLM call, finalize_experiment_X_Y.py rewrites the JSON without rerunning the paid model, preserving the expensive API results while fixing formatting or schema issues.

Practical Usage Examples

Running Exploratory Mode

To iterate on experimental parameters without generating official evidence:


# List all available flags

python experiment.py --help

# Run full RL vs LLM comparison with custom parameters

python experiment.py --mode both --model kimi-k3 --rl-episodes 5000

Executing the Canonical Protocol

To generate acceptance-grade evidence for Experiment 7-2:

python run_experiment_7_2.py

Internally, this script invokes:


# Simplified excerpt from run_experiment_7_2.py

from experiment import run_full_comparison

run_full_comparison(
    rl_episodes=10000,
    llm_episodes=20,
    model="kimi-k3",
    output_dir="validation/20260730_011704"
)

Finalizing Evidence Without API Costs

If the canonical run succeeded but JSON serialization failed:

python finalize_experiment_7_2.py validation/20260730_011704

This command validates the existing raw data and writes a compliant evidence.json suitable for book citation.

Summary

  • Experiments are modular: Each lives in its own directory under chapter<N>/<topic>/ with isolated agents, environments, and validation data.
  • Three-script architecture: experiment.py supports exploration, run_experiment_X_Y.py enforces canonical protocols, and finalize_experiment_X_Y.py handles post-hoc evidence formatting.
  • Versioned evidence: All official results serialize to validation/<timestamp>/evidence.json, creating an immutable audit trail for every figure in the book.
  • Numeric mapping: Script names like run_experiment_9_1.py correspond directly to experiment numbers cited in the book text (e.g., Chapter 9's trajectory verifier).

Frequently Asked Questions

What is the difference between experiment.py and run_experiment_X_Y.py?

experiment.py is the flexible CLI used during development to iterate on parameters and debug agents, while run_experiment_X_Y.py is the canonical script that executes a fixed protocol to generate citable evidence. The canonical runner enforces strict error handling and writes to validation/ directories, whereas the exploratory runner prioritizes convenience and configurability.

Where is the official evidence data stored?

Official evidence resides in validation/<timestamp>/evidence.json within each experiment directory. For example, chapter1/learning-from-experience/validation/20260730_011704/evidence.json contains the serialized raw LLM responses, timestamps, and metrics that the book cites as proof for Experiment 7-2.

How do experiment numbers map to the book's content?

The numeric identifiers in script names map directly to experiment numbers in the book text. For instance, run_experiment_7_2.py corresponds to Experiment 7-2 (RL vs LLM comparison in Chapter 1), while run_experiment_9_1.py handles Chapter 9's trajectory verifier experiment. The respective README.md files explicitly list which experiments are implemented in each directory.

Can I run experiments without incurring LLM API costs?

Yes, for development and testing, you can use experiment.py in Q-learning only mode (--mode rl-only) to validate environment logic and agent implementations without calling paid LLM APIs. Additionally, finalize_experiment_X_Y.py allows you to regenerate evidence JSON from existing raw data without making new API calls.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →