How Experiments Are Organized in the AI Agent Book: A Technical Guide
Experiments in the AI Agent Book follow a strict modular structure with three distinct scripts—experiment.py for exploration, run_experiment_X_Y.py for canonical evidence generation, and finalize_experiment_X_Y.py for post-processing—organized under chapter-specific directories with versioned validation artifacts.
The bojieli/ai-agent-book repository implements a rigorous, reproducible research framework where every experiment is encapsulated as a self-contained module. Understanding how experiments are organized in the AI Agent Book reveals a deliberate architecture that separates exploratory iteration from acceptance-grade evidence collection, ensuring every figure and table in the text can be traced to serialized raw data.
The Three-Tier Experiment Architecture
Each experiment directory contains a standardized trio of Python scripts that serve distinct phases of the research workflow. This separation prevents exploratory noise from contaminating citable results while allowing rapid iteration during development.
Exploratory Runner (experiment.py)
The experiment.py file serves as the flexible CLI entry point for iterative development. Located at paths like chapter1/learning-from-experience/experiment.py, this script supports dynamic parameter switching between Q-learning and LLM agents, custom episode counts, and interactive plotting capabilities.
Researchers use this script to explore hyperparameters, generate debugging plots, or compare model behaviors without strict protocol enforcement. The CLI accepts flags such as --mode (accepting values like both, rl-only, or llm-only), --rl-episodes, --llm-episodes, and --model (e.g., kimi-k3).
Canonical Runner (run_experiment_X_Y.py)
For deterministic, book-citable results, the repository provides canonical runners such as run_experiment_7_2.py and run_experiment_9_1.py. These scripts enforce strict protocols: fixed episode counts (e.g., 10,000 RL episodes and 20 LLM episodes for Experiment 7-2), mandatory raw response logging, and abort-on-error handling to ensure every run produces valid evidence.
The canonical runner calls underlying functions like run_full_comparison() with hardcoded parameters and writes timestamped output to validation/<timestamp>/evidence.json. Any deviation—such as missing response IDs, fallback parsers triggering, or API errors—causes immediate termination to preserve evidence integrity.
Evidence Finalizer (finalize_experiment_X_Y.py)
When serialization fails after successful LLM API calls, finalize_experiment_7_2.py regenerates the evidence JSON without incurring additional API costs. This script loads existing raw campaign data from the validation directory, validates it against the required schema, and writes a clean evidence.json that the book references directly.
Directory Structure and File Organization
Experiments are nested within chapter directories following a predictable schema. Each chapter lives under chapter<N>/ and contains topic-specific experiment directories (e.g., learning-from-experience, context-compression, trajectory-verifier).
The typical layout for an experiment directory includes:
chapter1/learning-from-experience/
├── experiment.py # Flexible CLI runner
├── run_experiment_7_2.py # Canonical runner (Experiment 7-2)
├── finalize_experiment_7_2.py # Evidence-only finalizer
├── env.example # API key template
├── game_environment.py # Environment logic
├── rl_agent.py # Q-learning implementation
├── llm_agent.py # LLM agent implementation
├── tests/ # Unit and integration tests
│ ├── test_basic.py
│ └── ...
├── validation/ # Serialized evidence
│ └── 20260730_011704/
│ └── evidence.json
└── README.md # Human overview (lists experiments 7-1, 7-2)
The README.md in each experiment directory provides human-readable context. For example, the Chapter 1 README explicitly lists Experiment 7-1 (Q-learning only) and Experiment 7-2 (full RL vs LLM comparison), along with quick-start commands for replication.
The Experiment Lifecycle
The organization of files supports a four-phase workflow that isolates exploration from publication-grade evidence generation:
-
Exploratory Phase – Developers run
experiment.pywith custom flags (e.g.,--mode both --model kimi-k3) to explore hyperparameters, generate plots, or debug agents without protocol constraints. -
Canonical Phase – When the experimental protocol is finalized, running
run_experiment_X_Y.pyexecutes the exact episode counts and logging requirements specified in the book text. This generates the official evidence used for citations. -
Evidence Generation – The canonical script writes a timestamped
evidence.jsonundervalidation/. This file contains serialized raw LLM responses, timestamps, and metrics that serve as the formal proof for the experiment. -
Post-Processing – If only the serialization step fails after a successful LLM call,
finalize_experiment_X_Y.pyrewrites the JSON without rerunning the paid model, preserving the expensive API results while fixing formatting or schema issues.
Practical Usage Examples
Running Exploratory Mode
To iterate on experimental parameters without generating official evidence:
# List all available flags
python experiment.py --help
# Run full RL vs LLM comparison with custom parameters
python experiment.py --mode both --model kimi-k3 --rl-episodes 5000
Executing the Canonical Protocol
To generate acceptance-grade evidence for Experiment 7-2:
python run_experiment_7_2.py
Internally, this script invokes:
# Simplified excerpt from run_experiment_7_2.py
from experiment import run_full_comparison
run_full_comparison(
rl_episodes=10000,
llm_episodes=20,
model="kimi-k3",
output_dir="validation/20260730_011704"
)
Finalizing Evidence Without API Costs
If the canonical run succeeded but JSON serialization failed:
python finalize_experiment_7_2.py validation/20260730_011704
This command validates the existing raw data and writes a compliant evidence.json suitable for book citation.
Summary
- Experiments are modular: Each lives in its own directory under
chapter<N>/<topic>/with isolated agents, environments, and validation data. - Three-script architecture:
experiment.pysupports exploration,run_experiment_X_Y.pyenforces canonical protocols, andfinalize_experiment_X_Y.pyhandles post-hoc evidence formatting. - Versioned evidence: All official results serialize to
validation/<timestamp>/evidence.json, creating an immutable audit trail for every figure in the book. - Numeric mapping: Script names like
run_experiment_9_1.pycorrespond directly to experiment numbers cited in the book text (e.g., Chapter 9's trajectory verifier).
Frequently Asked Questions
What is the difference between experiment.py and run_experiment_X_Y.py?
experiment.py is the flexible CLI used during development to iterate on parameters and debug agents, while run_experiment_X_Y.py is the canonical script that executes a fixed protocol to generate citable evidence. The canonical runner enforces strict error handling and writes to validation/ directories, whereas the exploratory runner prioritizes convenience and configurability.
Where is the official evidence data stored?
Official evidence resides in validation/<timestamp>/evidence.json within each experiment directory. For example, chapter1/learning-from-experience/validation/20260730_011704/evidence.json contains the serialized raw LLM responses, timestamps, and metrics that the book cites as proof for Experiment 7-2.
How do experiment numbers map to the book's content?
The numeric identifiers in script names map directly to experiment numbers in the book text. For instance, run_experiment_7_2.py corresponds to Experiment 7-2 (RL vs LLM comparison in Chapter 1), while run_experiment_9_1.py handles Chapter 9's trajectory verifier experiment. The respective README.md files explicitly list which experiments are implemented in each directory.
Can I run experiments without incurring LLM API costs?
Yes, for development and testing, you can use experiment.py in Q-learning only mode (--mode rl-only) to validate environment logic and agent implementations without calling paid LLM APIs. Additionally, finalize_experiment_X_Y.py allows you to regenerate evidence JSON from existing raw data without making new API calls.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →