# How the Agent Manager Orchestrates Experiments in AI Scientist v2

> Discover how the Agent Manager orchestrates experiments in AI Scientist v2. It manages research stages, spawns agents, and drives experiments autonomously based on LLM evaluations.

- Repository: [Sakana AI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2)
- Tags: internals
- Published: 2026-03-28

---

**The Agent Manager serves as the central conductor of AI Scientist v2's research loop, hierarchically managing four main stages through dynamically generated sub-stages, spawning ParallelAgents for execution, and autonomously advancing experiments based on LLM-evaluated completion criteria.**

The **Agent Manager** is the core orchestration engine in SakanaAI/AI-Scientist-v2 that transforms a high-level research task into an autonomous experimental workflow. According to the source code in [`ai_scientist/treesearch/agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/agent_manager.py), it maintains the complete experiment state—including task descriptions, configuration, workspace directories, and stage histories—while coordinating between planning, execution, and evaluation components.

## Architecture of the Orchestration System

The Agent Manager does not execute experiments directly. Instead, it delegates computational work to specialized components while retaining decision-making authority over stage progression.

### Core Components

| Component | Responsibility | Source Location |
|-----------|----------------|-----------------|
| `AgentManager` | Holds experiment state, creates stages, spawns agents, evaluates completion, and manages transitions | [`ai_scientist/treesearch/agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/agent_manager.py) |
| `ParallelAgent` | Executes LLM-generated plans and code in isolated processes, gathers metrics and feedback | [`ai_scientist/treesearch/parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/parallel_agent.py) |
| `Journal` / `Node` | Persistent per-stage logs storing trials, code, results, and metrics with helper methods like `get_best_node` | [`ai_scientist/treesearch/journal.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/journal.py) |
| `Backend.query` | LLM function-calling interface for structured responses (stage configs, evaluations) | [`ai_scientist/treesearch/backend.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/backend.py) |

### The Four Main Research Stages

The manager structures experiments into a fixed pipeline of **main stages**:

1. `initial_implementation` (Stage 1)
2. `baseline_tuning` (Stage 2)
3. `creative_research` (Stage 3)
4. `ablation_studies` (Stage 4)

Each main stage contains one or more **sub-stages** that are generated dynamically based on intermediate results. The `main_stage_dict` mapping (defined in [`agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/agent_manager.py) lines 43-48) encodes this progression, while `parse_stage_names` (lines 127-141) extracts identifiers from stage strings like `"2_baseline_tuning_3_second_attempt"` to determine appropriate configuration logic.

## Experiment Initialization

When instantiated, the Agent Manager validates the research task and prepares the experimental framework.

```python
def __init__(self, task_desc: str, cfg: Any, workspace_dir: Path):
    self.task_desc = json.loads(task_desc)          # validated JSON task description

    self.cfg = cfg
    self.workspace_dir = workspace_dir
    self.current_stage_number = 0
    self.stages: List[Stage] = []
    self.journals: Dict[str, Journal] = {}
    self.stage_history: List[StageTransition] = []
    self.main_stage_dict = {1: "initial_implementation", 2: "baseline_tuning",
                           3: "creative_research", 4: "ablation_studies"}
    self._create_initial_stage()                    # first stage ready

```

The constructor performs three critical setup operations:

- **Task validation**: Parses and stores the JSON task description (lines 24-34)
- **State initialization**: Creates empty containers for stages, journals, and transition history
- **Stage template loading**: Builds the four high-level research stage definitions
- **Initial stage creation**: Calls `_create_initial_stage()` to bootstrap the first sub-stage

## The Orchestration Loop

The `run` method (lines 92-130) implements the nested loop structure that drives the entire experiment. This is where the Agent Manager orchestrates the iterative research process.

```python
def run(self, exec_callback, step_callback=None):
    while self.current_stage:                      # Main-stage loop

        main_stage = self.parse_stage_names(self.current_stage.name)[0]
        current_substage = self.current_stage
        while current_substage:                    # Sub-stage loop

            with self._create_agent_for_stage(current_substage) as agent:
                # Feed best node from previous stage

                if self.stage_history:
                    prev_stage = self.stage_history[-1].from_stage
                    prev_best = self._get_best_implementation(prev_stage)
                    if prev_best:
                        self.journals[self.current_stage.name].append(prev_best)

                # Iterative stepping until completion

                while True:
                    agent.step(exec_callback)      # ParallelAgent runs trial

                    if step_callback:
                        step_callback(current_substage,
                                      self.journals[current_substage.name])

                    # Check main stage completion

                    main_stage_complete, main_stage_feedback = \
                        self._check_stage_completion(current_substage)

                    # Check sub-stage completion

                    substage_complete, substage_feedback = \
                        self._check_substage_completion(current_substage,
                                                        self.journals[current_substage.name])

```

### Agent Creation and Context Passing

For each sub-stage, `_create_agent_for_stage` (lines 74-86) instantiates a **ParallelAgent** with stage-specific configuration. Crucially, if previous stages exist, the manager retrieves the best implementation node via `_get_best_implementation` and appends it to the current journal. This ensures experimental continuity—new agents start from the latest validated code rather than from scratch.

### Execution Callback Integration

The `exec_callback` parameter represents the sandbox boundary. While the manager orchestrates *what* to run, the callback handles *how* to run it—typically executing generated code in isolated Docker containers or subprocesses. The `agent.step()` call triggers the full trial cycle: plan generation, code synthesis, execution via the callback, and metric extraction.

## Stage Completion and Transitions

The Agent Manager employs a hierarchical evaluation strategy to determine when to advance experiments.

### Sub-Stage Completion

`_check_substage_completion` (lines 44-88) evaluates whether a specific sub-stage's goals have been satisfied. It uses the `stage_completion_eval_spec` function specification (defined in [`backend.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/backend.py)) to prompt the LLM for structured evaluation of the current journal contents. If the sub-stage completes, the manager calls `_create_next_substage` to generate fresh goals based on current metrics and identified issues.

### Main Stage Completion

`_check_stage_completion` (lines 110-165) operates at a higher granularity, examining:
- Overall journal size and node counts
- Number of "good" nodes (successful implementations)
- Stage-specific criteria (e.g., for stage 2, validating that the best node improves over the baseline)

When a main stage completes, `_create_next_main_stage` (lines 166-189) advances to the next entry in `main_stage_dict`, creating the first sub-stage of the subsequent research phase.

### Stage History Tracking

Every transition is recorded as a `StageTransition` object in `self.stage_history`, maintaining a complete audit trail of the experimental progression from implementation through ablation studies.

## Checkpointing and Persistence

To enable resumable long-running research, the manager implements robust state serialization.

`_save_checkpoint` (lines 49-73) pickles the entire experiment state—including all stages, journals, and history—after each main stage completion. This allows researchers to resume complex multi-day experiments without losing progress or intermediate results.

## Running an Experiment: Practical Example

### Basic Experiment Execution

```python
from pathlib import Path
from ai_scientist.treesearch.agent_manager import AgentManager
from ai_scientist.utils.config import Config

# Load task description

with open("task_description.json", "r") as f:
    task_json = f.read()

# Load configuration (model names, workers, timeouts)

cfg = Config.from_yaml("config.yaml")

# Prepare workspace

workspace = Path("./workspace/run_001")

# Initialize the orchestrator

manager = AgentManager(task_desc=task_json, cfg=cfg, workspace_dir=workspace)

# Define execution sandbox (simplified example)

def exec_callback(code: str) -> dict:
    # In production, this runs code in isolated containers

    return {"status": "executed", "metrics": {}}

# Run the full research pipeline

manager.run(exec_callback)

```

### Inspecting Experimental Results

After execution or checkpoint recovery, access the experimental record through the journal interface:

```python

# Access specific sub-stage journal

journal = manager.journals["2_baseline_tuning_1_first_attempt"]

# Retrieve best implementation

best_node = journal.get_best_node(cfg=manager.cfg)
print(f"Best metric: {best_node.metric.value}")
print(f"Implementation:\n{best_node.code}")

```

### Manual Stage Advancement

For debugging or custom workflows, force sub-stage creation manually:

```python
current = manager.current_stage
journal = manager.journals[current.name]

# Create next sub-stage with custom feedback

next_sub = manager._create_next_substage(
    current, 
    journal, 
    substage_feedback="manual advancement"
)

if next_sub:
    manager.stages.append(next_sub)
    manager.journals[next_sub.name] = Journal()
    manager.current_stage = next_sub

```

## Summary

- The **Agent Manager** in [`ai_scientist/treesearch/agent_manager.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/treesearch/agent_manager.py) implements a hierarchical stage system (main stages 1-4 containing dynamic sub-stages) to structure autonomous research.
- **ParallelAgent** instances handle actual code execution in isolated processes, while the manager maintains state through **Journal** and **Node** objects.
- The `run` method implements a nested loop architecture: main stage loop → sub-stage loop → iterative agent stepping.
- Completion is determined by LLM-evaluated criteria via `_check_substage_completion` and `_check_stage_completion`, enabling adaptive progression based on experimental results.
- The system supports full persistence through pickle-based checkpointing after each main stage, ensuring recoverability for long-running experiments.

## Frequently Asked Questions

### How does the Agent Manager decide when to move to the next research stage?

The manager employs a two-level evaluation system. First, `_check_substage_completion` queries the LLM to evaluate if current sub-stage goals are met. If satisfied, it either creates a new sub-stage with refined goals or, if sufficient successful trials exist, `_check_stage_completion` validates main-stage criteria (such as metric improvement over baselines). Only when both evaluations pass does `_create_next_main_stage` advance to the next phase of the research pipeline.

### What happens if a generated experiment fails or produces poor results?

Failed trials are stored as **Node** objects in the stage **Journal** with their associated metrics and execution logs. The manager's `_create_next_substage` method analyzes these failures via LLM evaluation to generate improved goals for subsequent attempts. Additionally, `get_best_node` helpers allow the system to propagate only the highest-performing implementations forward to new stages, effectively filtering out unsuccessful experiments while learning from them.

### Can experiments be resumed after interruption?

Yes. The `_save_checkpoint` method serializes the complete experiment state—including all stages, journals, and transition histories—to a pickle file after each main stage completion. Researchers can reload the `AgentManager` from this checkpoint to resume exactly where the experiment left off, preserving all intermediate implementations and metrics without requiring restart from stage one.

### How does the Agent Manager handle code execution safety?

The manager itself does not execute code directly. Instead, it delegates execution to the `exec_callback` function passed to `run()`. This architectural boundary allows users to inject arbitrary sandboxing—such as Docker containers, firejail, or restricted subprocesses—between the orchestration logic and generated code execution. The [`parallel_agent.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/parallel_agent.py) module handles the execution mechanics, but the safety policy is determined by the callback implementation.