How the Agent Manager Orchestrates Experiments in AI Scientist v2
The Agent Manager serves as the central conductor of AI Scientist v2's research loop, hierarchically managing four main stages through dynamically generated sub-stages, spawning ParallelAgents for execution, and autonomously advancing experiments based on LLM-evaluated completion criteria.
The Agent Manager is the core orchestration engine in SakanaAI/AI-Scientist-v2 that transforms a high-level research task into an autonomous experimental workflow. According to the source code in ai_scientist/treesearch/agent_manager.py, it maintains the complete experiment state—including task descriptions, configuration, workspace directories, and stage histories—while coordinating between planning, execution, and evaluation components.
Architecture of the Orchestration System
The Agent Manager does not execute experiments directly. Instead, it delegates computational work to specialized components while retaining decision-making authority over stage progression.
Core Components
| Component | Responsibility | Source Location |
|---|---|---|
AgentManager |
Holds experiment state, creates stages, spawns agents, evaluates completion, and manages transitions | ai_scientist/treesearch/agent_manager.py |
ParallelAgent |
Executes LLM-generated plans and code in isolated processes, gathers metrics and feedback | ai_scientist/treesearch/parallel_agent.py |
Journal / Node |
Persistent per-stage logs storing trials, code, results, and metrics with helper methods like get_best_node |
ai_scientist/treesearch/journal.py |
Backend.query |
LLM function-calling interface for structured responses (stage configs, evaluations) | ai_scientist/treesearch/backend.py |
The Four Main Research Stages
The manager structures experiments into a fixed pipeline of main stages:
initial_implementation(Stage 1)baseline_tuning(Stage 2)creative_research(Stage 3)ablation_studies(Stage 4)
Each main stage contains one or more sub-stages that are generated dynamically based on intermediate results. The main_stage_dict mapping (defined in agent_manager.py lines 43-48) encodes this progression, while parse_stage_names (lines 127-141) extracts identifiers from stage strings like "2_baseline_tuning_3_second_attempt" to determine appropriate configuration logic.
Experiment Initialization
When instantiated, the Agent Manager validates the research task and prepares the experimental framework.
def __init__(self, task_desc: str, cfg: Any, workspace_dir: Path):
self.task_desc = json.loads(task_desc) # validated JSON task description
self.cfg = cfg
self.workspace_dir = workspace_dir
self.current_stage_number = 0
self.stages: List[Stage] = []
self.journals: Dict[str, Journal] = {}
self.stage_history: List[StageTransition] = []
self.main_stage_dict = {1: "initial_implementation", 2: "baseline_tuning",
3: "creative_research", 4: "ablation_studies"}
self._create_initial_stage() # first stage ready
The constructor performs three critical setup operations:
- Task validation: Parses and stores the JSON task description (lines 24-34)
- State initialization: Creates empty containers for stages, journals, and transition history
- Stage template loading: Builds the four high-level research stage definitions
- Initial stage creation: Calls
_create_initial_stage()to bootstrap the first sub-stage
The Orchestration Loop
The run method (lines 92-130) implements the nested loop structure that drives the entire experiment. This is where the Agent Manager orchestrates the iterative research process.
def run(self, exec_callback, step_callback=None):
while self.current_stage: # Main-stage loop
main_stage = self.parse_stage_names(self.current_stage.name)[0]
current_substage = self.current_stage
while current_substage: # Sub-stage loop
with self._create_agent_for_stage(current_substage) as agent:
# Feed best node from previous stage
if self.stage_history:
prev_stage = self.stage_history[-1].from_stage
prev_best = self._get_best_implementation(prev_stage)
if prev_best:
self.journals[self.current_stage.name].append(prev_best)
# Iterative stepping until completion
while True:
agent.step(exec_callback) # ParallelAgent runs trial
if step_callback:
step_callback(current_substage,
self.journals[current_substage.name])
# Check main stage completion
main_stage_complete, main_stage_feedback = \
self._check_stage_completion(current_substage)
# Check sub-stage completion
substage_complete, substage_feedback = \
self._check_substage_completion(current_substage,
self.journals[current_substage.name])
Agent Creation and Context Passing
For each sub-stage, _create_agent_for_stage (lines 74-86) instantiates a ParallelAgent with stage-specific configuration. Crucially, if previous stages exist, the manager retrieves the best implementation node via _get_best_implementation and appends it to the current journal. This ensures experimental continuity—new agents start from the latest validated code rather than from scratch.
Execution Callback Integration
The exec_callback parameter represents the sandbox boundary. While the manager orchestrates what to run, the callback handles how to run it—typically executing generated code in isolated Docker containers or subprocesses. The agent.step() call triggers the full trial cycle: plan generation, code synthesis, execution via the callback, and metric extraction.
Stage Completion and Transitions
The Agent Manager employs a hierarchical evaluation strategy to determine when to advance experiments.
Sub-Stage Completion
_check_substage_completion (lines 44-88) evaluates whether a specific sub-stage's goals have been satisfied. It uses the stage_completion_eval_spec function specification (defined in backend.py) to prompt the LLM for structured evaluation of the current journal contents. If the sub-stage completes, the manager calls _create_next_substage to generate fresh goals based on current metrics and identified issues.
Main Stage Completion
_check_stage_completion (lines 110-165) operates at a higher granularity, examining:
- Overall journal size and node counts
- Number of "good" nodes (successful implementations)
- Stage-specific criteria (e.g., for stage 2, validating that the best node improves over the baseline)
When a main stage completes, _create_next_main_stage (lines 166-189) advances to the next entry in main_stage_dict, creating the first sub-stage of the subsequent research phase.
Stage History Tracking
Every transition is recorded as a StageTransition object in self.stage_history, maintaining a complete audit trail of the experimental progression from implementation through ablation studies.
Checkpointing and Persistence
To enable resumable long-running research, the manager implements robust state serialization.
_save_checkpoint (lines 49-73) pickles the entire experiment state—including all stages, journals, and history—after each main stage completion. This allows researchers to resume complex multi-day experiments without losing progress or intermediate results.
Running an Experiment: Practical Example
Basic Experiment Execution
from pathlib import Path
from ai_scientist.treesearch.agent_manager import AgentManager
from ai_scientist.utils.config import Config
# Load task description
with open("task_description.json", "r") as f:
task_json = f.read()
# Load configuration (model names, workers, timeouts)
cfg = Config.from_yaml("config.yaml")
# Prepare workspace
workspace = Path("./workspace/run_001")
# Initialize the orchestrator
manager = AgentManager(task_desc=task_json, cfg=cfg, workspace_dir=workspace)
# Define execution sandbox (simplified example)
def exec_callback(code: str) -> dict:
# In production, this runs code in isolated containers
return {"status": "executed", "metrics": {}}
# Run the full research pipeline
manager.run(exec_callback)
Inspecting Experimental Results
After execution or checkpoint recovery, access the experimental record through the journal interface:
# Access specific sub-stage journal
journal = manager.journals["2_baseline_tuning_1_first_attempt"]
# Retrieve best implementation
best_node = journal.get_best_node(cfg=manager.cfg)
print(f"Best metric: {best_node.metric.value}")
print(f"Implementation:\n{best_node.code}")
Manual Stage Advancement
For debugging or custom workflows, force sub-stage creation manually:
current = manager.current_stage
journal = manager.journals[current.name]
# Create next sub-stage with custom feedback
next_sub = manager._create_next_substage(
current,
journal,
substage_feedback="manual advancement"
)
if next_sub:
manager.stages.append(next_sub)
manager.journals[next_sub.name] = Journal()
manager.current_stage = next_sub
Summary
- The Agent Manager in
ai_scientist/treesearch/agent_manager.pyimplements a hierarchical stage system (main stages 1-4 containing dynamic sub-stages) to structure autonomous research. - ParallelAgent instances handle actual code execution in isolated processes, while the manager maintains state through Journal and Node objects.
- The
runmethod implements a nested loop architecture: main stage loop → sub-stage loop → iterative agent stepping. - Completion is determined by LLM-evaluated criteria via
_check_substage_completionand_check_stage_completion, enabling adaptive progression based on experimental results. - The system supports full persistence through pickle-based checkpointing after each main stage, ensuring recoverability for long-running experiments.
Frequently Asked Questions
How does the Agent Manager decide when to move to the next research stage?
The manager employs a two-level evaluation system. First, _check_substage_completion queries the LLM to evaluate if current sub-stage goals are met. If satisfied, it either creates a new sub-stage with refined goals or, if sufficient successful trials exist, _check_stage_completion validates main-stage criteria (such as metric improvement over baselines). Only when both evaluations pass does _create_next_main_stage advance to the next phase of the research pipeline.
What happens if a generated experiment fails or produces poor results?
Failed trials are stored as Node objects in the stage Journal with their associated metrics and execution logs. The manager's _create_next_substage method analyzes these failures via LLM evaluation to generate improved goals for subsequent attempts. Additionally, get_best_node helpers allow the system to propagate only the highest-performing implementations forward to new stages, effectively filtering out unsuccessful experiments while learning from them.
Can experiments be resumed after interruption?
Yes. The _save_checkpoint method serializes the complete experiment state—including all stages, journals, and transition histories—to a pickle file after each main stage completion. Researchers can reload the AgentManager from this checkpoint to resume exactly where the experiment left off, preserving all intermediate implementations and metrics without requiring restart from stage one.
How does the Agent Manager handle code execution safety?
The manager itself does not execute code directly. Instead, it delegates execution to the exec_callback function passed to run(). This architectural boundary allows users to inject arbitrary sandboxing—such as Docker containers, firejail, or restricted subprocesses—between the orchestration logic and generated code execution. The parallel_agent.py module handles the execution mechanics, but the safety policy is determined by the callback implementation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →