# What Datasets Are Used in Stage 3 of AI Scientist v2: Multi-Dataset Evaluation Explained

> Discover the datasets used in Stage 3 of AI Scientist v2. Learn how standard benchmarks and synthetic variants are evaluated for robustness and visualized.

- Repository: [Sakana AI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2)
- Tags: deep-dive
- Published: 2026-03-28

---

**Stage 3 of AI Scientist v2 evaluates experiments on multiple datasets—including standard benchmarks and user-specified synthetic variants—to generate comparative visualizations and validate robustness across data variations.**

Stage 3 of the AI Scientist v2 pipeline, developed by SakanaAI, serves as the *"baseline tuning & dataset-wide comparison"* phase. This stage processes the **datasets used in Stage 3 of AI Scientist v2** that are inherited from previous execution phases, requiring successful validation across multiple data sources to ensure comprehensive experimental verification. The system leverages these datasets to produce cross-dataset comparison plots and confirm research validity through automated vision-language model analysis.

## The Role of Stage 3 in the AI Scientist v2 Pipeline

Stage 3 functions as the robustness validation layer of the AI Scientist v2 workflow. During this phase, the system expects the experiment to have been executed on **multiple datasets** rather than a single source. These datasets can include standard benchmark collections or synthetic variants generated according to user specifications.

The primary objective is to move beyond single-dataset validation and establish whether the proposed scientific approach generalizes across different data distributions. This multi-dataset requirement ensures that the generated research findings are not artifacts of a specific data sample but represent genuine algorithmic improvements.

## Dataset Sources and Configuration

### Inherited Data Structures from Stage 2

The datasets entering Stage 3 originate from the execution phase completed in Stage 2. According to the implementation in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) (lines 886-902), the plotting code for Stage 3 initializes with the code generated during Stage 2, which stores results for each dataset in a NumPy dictionary called `experiment_data`.

The LLM prompt explicitly instructs the model to maintain this data structure while extending it for comparative analysis:

```python

# Prompt fragment from parallel_agent.py guiding LLM dataset handling

prompt_guideline.extend([
    "IMPORTANT: Use the following base plotting code as a starting point:",
    "Base plotting code: " + plot_code_from_prev_stage,
    "Modify the base plotting code to:",
    "1. Keep the same numpy data structure and plotting style",
    "2. Add comparison plots between different datasets",
    "3. Add dataset-specific visualizations if needed",
    "4. Include clear labels indicating which plots are from which dataset",
    "5. Use consistent naming conventions for saved files",
])

```

### Configuring the Number of Synthetic Datasets

The system allows explicit configuration of how many synthetic datasets must be evaluated. In [`bfts_config.yaml`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/bfts_config.yaml) (lines 30-33), the parameter `num_syn_datasets` defaults to `1`, but users can increase this value to force evaluation across multiple synthetic variants.

When `cfg.experiment.num_syn_datasets` is set greater than 1, the agent in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) (lines 315-326) explicitly instructs the experiment to evaluate on *at least* that many different synthetic datasets and report separate metrics for each. This configuration ensures the system generates robustness analyses across data variations rather than single-point validations.

## Dataset Verification and Completion Logic

### VLM Analysis and Success Tracking

After plot generation, a **Vision-Language Model (VLM)** parses the visual analyses to determine which datasets produced valid experimental results. The code in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) (lines 835-872) implements logic that extracts the list of successfully tested datasets and stores it in `node.datasets_successfully_tested`.

This verification step ensures that only datasets producing coherent, analyzable results count toward Stage 3 completion:

```python

# Extracting validated datasets from VLM analysis

datasets_successfully_tested = node.datasets_successfully_tested
if not datasets_successfully_tested:
    print("No datasets passed the VLM check yet.")
else:
    print("Datasets confirmed in Stage 3:", datasets_successfully_tested)

```

### Stage 3 Completion Criteria

According to [`ai_scientist/treesearch/agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/agent_manager.py) (lines 500-530), Stage 3 completion requires three specific conditions related to dataset validation:

1. **Result diversity**: The best node must differ from the initial node, indicating meaningful experimental evolution
2. **Dataset validity**: The node must contain `datasets_successfully_tested` with valid results from multiple sources
3. **Execution legitimacy**: The experiment runtime must exceed a minimum threshold to avoid trivial or cached results

The manager specifically checks that `node.datasets_successfully_tested` contains entries before marking the stage as complete, ensuring that multi-dataset evaluation has actually occurred.

## Summary

- Stage 3 evaluates experiments across **multiple datasets** inherited from Stage 2, stored in the `experiment_data` NumPy structure.
- Users configure the required number of synthetic datasets via `num_syn_datasets` in [`bfts_config.yaml`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/bfts_config.yaml), with the system enforcing evaluation on at least that many variants.
- A **Vision-Language Model** verifies dataset execution success, populating `node.datasets_successfully_tested` with validated data sources.
- Completion requires both algorithmic improvement (best node ≠ initial node) and confirmed multi-dataset results with reasonable execution times.

## Frequently Asked Questions

### What types of datasets are supported in Stage 3 of AI Scientist v2?

Stage 3 supports both standard benchmark datasets and synthetic variants specified by the user. The system is dataset-agnostic regarding content but requires multiple distinct sources to generate comparative analyses. According to the source code in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py), these datasets are carried forward from Stage 2 execution and must be present in the `experiment_data` dictionary structure.

### How does AI Scientist v2 verify that datasets executed correctly in Stage 3?

The system employs a Vision-Language Model (VLM) to parse generated plots and analyses, as implemented in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) (lines 835-872). This VLM validation produces the `datasets_successfully_tested` list, which identifies which specific datasets produced valid, analyzable results. Only datasets appearing in this list count toward Stage 3 completion criteria.

### Can I configure how many datasets Stage 3 requires?

Yes. The [`bfts_config.yaml`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/bfts_config.yaml) file (lines 30-33) contains the `num_syn_datasets` parameter, which defaults to 1 but can be increased to require evaluation across multiple synthetic datasets. When this value exceeds 1, the agent explicitly instructs the LLM to evaluate on at least that many synthetic variants and report metrics separately for each, as seen in [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) (lines 315-326).

### What happens if no datasets pass the VLM verification in Stage 3?

If `node.datasets_successfully_tested` remains empty after VLM analysis, Stage 3 cannot complete successfully. According to the completion logic in [`ai_scientist/treesearch/agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/agent_manager.py) (lines 500-530), the manager requires valid dataset results to mark the stage finished. The system will continue attempting to generate valid experimental results until at least one dataset passes the VLM check and the best node differs from the initial node.