The Four Stages of Experiment Execution in AI Scientist v2: A Complete Technical Guide
AI Scientist v2 orchestrates automated research through four distinct stages—initial implementation, baseline tuning, creative research, and ablation studies—managed by the AgentManager class to systematically evolve experiments from prototype to publication-ready analysis.
The SakanaAI/AI-Scientist-v2 repository implements a structured pipeline for autonomous scientific discovery. Understanding the stages of experiment execution in AI Scientist v2 is essential for researchers configuring automated literature searches, as each phase builds upon the previous with specific constraints regarding datasets, hyperparameters, and architectural modifications.
Overview of the Experiment Execution Pipeline
AI Scientist v2 breaks down complex research tasks into a four-stage roadmap defined in ai_scientist/treesearch/agent_manager.py. The AgentManager class maintains two critical dictionaries that govern the entire lifecycle:
main_stage_dict: Maps stage numbers (1-4) to internal stage namesmain_stage_goals: Defines the objectives and constraints for each phase
The system uses a Stage dataclass (containing fields for name, description, goals, max_iterations, num_drafts, and stage_number) to instantiate each phase dynamically during runtime.
The Four Stages of AI Scientist v2 Experiment Execution
Stage 1: Initial Implementation
The initial_implementation stage focuses on establishing functional correctness. According to the source code in agent_manager.py, the goals for this phase are:
- Build a basic, working prototype using a simple dataset
- Aim for basic functional correctness rather than optimal performance
- Leverage provided "Code To Use" as a starting point when available
This stage creates the foundation through AgentManager._create_initial_stage(), which generates a Stage instance named "1_initial_implementation_1_preliminary".
Stage 2: Baseline Tuning
The baseline_tuning stage strictly prohibits architectural changes while optimizing hyperparameters. The main_stage_goals dictionary specifies:
- Modify hyperparameters such as learning rate, number of epochs, and batch size
- Do not change the model architecture from the previous stage
- Introduce two additional Hugging Face datasets for expanded testing
This constraint ensures that performance improvements stem solely from optimization rather than structural modifications.
Stage 3: Creative Research
The creative_research stage removes architectural constraints to encourage novel exploration. The goals emphasize:
- Exploring novel improvements and creative experiments
- Revealing new insights through innovative approaches
- Using three Hugging Face datasets in total for comprehensive validation
This phase allows the LLM agent to "think outside the box" and propose substantive architectural or methodological innovations.
Stage 4: Ablation Studies
The final ablation_studies stage conducts systematic component-wise analysis. The requirements include:
- Performing component ablations to quantify each part's contribution
- Using the same datasets from the previous stage to ensure consistency
- Generating empirical evidence for the importance of specific architectural choices
This stage produces the rigorous analysis required for scientific publication.
How the AgentManager Orchestrates Stage Progression
Stage instantiation follows a hybrid approach combining hardcoded initialization with LLM-generated configuration. The AgentManager class in ai_scientist/treesearch/agent_manager.py defines the roadmap as follows:
# ai_scientist/treesearch/agent_manager.py
self.main_stage_dict = {
1: "initial_implementation",
2: "baseline_tuning",
3: "creative_research",
4: "ablation_studies",
}
self.main_stage_goals = {
1: """
- Focus on getting basic working implementation
- Use a simple dataset
- Aim for basic functional correctness
- If you are given "Code To Use", you can directly use it as a starting point.""",
2: """
- Change hyperparameters such as learning rate, number of epochs, batch size, etc.
- DO NOT change the model architecture from the previous stage
- Introduce TWO more new datasets from HuggingFace for testing.""",
3: """
- Explore novel improvements
- Come up with experiments to reveal new insights
- Be creative and think outside the box
- MAKE SURE you use THREE HuggingFace dataset in total to test your models""",
4: """
- Conduct systematic component analysis that reveals the contribution of each part
- Use the same datasets you used from the previous stage"""
}
The initial stage is created via _create_initial_stage(), while subsequent stages are generated by querying the LLM using the generate_stage_config function specification (referenced as stage_config_spec in the codebase). All stages are stored in self.stages for sequential execution.
Tracking and Visualizing Stage Execution
During the experiment loop defined in ai_scientist/treesearch/perform_experiments_bfts_with_agentmanager.py, the system persists detailed stage metadata through a callback mechanism:
# ai_scientist/treesearch/perform_experiments_bfts_with_agentmanager.py
def step_callback(stage, journal):
# … generate findings and best metric …
notes_dir = cfg.log_dir / f"stage_{stage.name}" / "notes"
notes_dir.mkdir(parents=True, exist_ok=True)
stage_summary = {
"stage": stage.name,
"total_nodes": len(journal.nodes),
"buggy_nodes": len(journal.buggy_nodes),
"good_nodes": len(journal.good_nodes),
"best_metric": str(best_metric.metric) if best_metric else "None",
"current_findings": current_findings,
}
with open(notes_dir / "stage_progress.json", "w") as f:
json.dump(stage_summary, f, indent=2)
# Persist the whole run for this stage
save_run(cfg, journal, stage_name=f"stage_{stage.name}")
For monitoring progress across the pipeline, ai_scientist/treesearch/utils/tree_export.py scans the log directory for stage_<n> folders and validates completion by checking for tree_data.json, tree_plot.html, and journal.json files:
# ai_scientist/treesearch/utils/tree_export.py
def get_completed_stages(log_dir):
completed = []
for stage_num in range(1, 5):
prefix = f"stage_{stage_num}"
matching_dirs = [d for d in log_dir.iterdir()
if d.name.startswith(prefix)]
for stage_dir in matching_dirs:
if (stage_dir / "tree_data.json").exists() and \
(stage_dir / "tree_plot.html").exists() and \
(stage_dir / "journal.json").exists():
completed.append(f"Stage_{stage_num}")
break
return completed
This utility generates a unified HTML tree visualization listing all completed stages, enabling researchers to audit the progression from initial prototype through final ablation analysis.
Summary
- AI Scientist v2 implements a rigid four-stage pipeline (initial implementation → baseline tuning → creative research → ablation studies) to structure autonomous research.
- The
AgentManagerclass inagent_manager.pydefines stage goals that enforce specific constraints regarding architectural changes and dataset usage. - Stage progression combines programmatic initialization for stage 1 with LLM-generated configurations for stages 2-4 using the
generate_stage_configspecification. - Execution logs are stored in
stage_<n>subdirectories within the configured log directory, containing JSON summaries of node counts, metrics, and findings. - Visualization tools in
tree_export.pyenable HTML-based inspection of completed stages by scanning for required data files.
Frequently Asked Questions
What is the purpose of the baseline tuning stage?
The baseline tuning stage optimizes model performance exclusively through hyperparameter adjustments (learning rate, epochs, batch size) without modifying the underlying architecture. According to the main_stage_goals dictionary in agent_manager.py, this stage also mandates adding two new Hugging Face datasets to establish a robust performance baseline before architectural experiments begin.
How does AI Scientist v2 transition between stages?
Stage transitions are managed by the AgentManager class, which initializes the first stage through _create_initial_stage() and generates subsequent stages by querying the LLM with the generate_stage_config function specification. Each stage is instantiated as a Stage dataclass object and appended to self.stages for sequential processing during the experiment loop.
Where are stage execution logs stored?
Per-stage execution data is written to subdirectories following the pattern stage_<name>/ within the configured log_dir. Each stage folder contains a notes/ subdirectory with stage_progress.json (tracking node counts, best metrics, and current findings) plus serialized journal data. The save_run() function in the configuration utilities handles persistence of complete run states.
Can the number of stages be customized?
While the default implementation enforces four stages defined in main_stage_dict, the architecture supports dynamic stage generation through the LLM configuration system. The Stage dataclass and generate_stage_config mechanism theoretically allow for extended pipelines, though the current main_stage_dict hardcodes stages 1-4 with specific research-phase goals.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →