AI Scientist v1 vs v2: How the Template-Driven System Evolved into an Agentic Research Engine

AI Scientist v1 relies on rigid human-written markdown templates to orchestrate research workflows, whereas AI Scientist v2 eliminates templates entirely in favor of a progressive agentic tree search driven by an Experiment Manager that dynamically proposes, tests, and refines scientific ideas.

The SakanaAI/AI-Scientist-v2 repository represents a fundamental architectural overhaul of the original AI Scientist framework. While AI Scientist v1 demanded handcrafted templates for every stage of the research pipeline, the v2 codebase implements a fully autonomous discovery system capable of tackling open-ended scientific questions without predefined structural constraints.

Core Architectural Differences Between AI Scientist v1 and v2

AI Scientist v1 operates as a template-driven engine. According to the repository documentation, it executes linear shell scripts that render static markdown templates for hypothesis generation, experiment design, and paper writing. This approach yields high success rates when the research problem aligns perfectly with the template's assumptions, but it fails on novel or speculative scientific questions that lack existing template patterns.

AI Scientist v2 rewrites the system from the ground up to be template-free and agentic. Instead of executing predetermined scripts, v2 employs a progressive agentic tree search coordinated by the AgentManager class in ai_scientist/treesearch/agent_manager.py. This manager repeatedly proposes implementations, evaluates results, and refines experiments without human-authored structural guidance. The README explicitly states: "Unlike its predecessor (AI Scientist‑v1), the AI Scientist‑v2 removes reliance on human‑authored templates, generalizes across Machine Learning (ML) domains, and employs a progressive agentic tree search, guided by an experiment manager agent."

Key distinctions include:

  • Workflow Orchestration: v1 uses linear execution of shell scripts; v2 uses dynamic multi-stage tree search with checkpointing.
  • Research Scope: v1 works only where templates exist; v2 generalizes across domains through automated discovery.
  • Core Components: v2 introduces ParallelAgents (ai_scientist/treesearch/parallel_agent.py) for concurrent execution, a VLM (vision-language model) for automated plot analysis, and a backend abstraction layer for LLM calls.
  • Extensibility: Adding a new domain to v1 requires writing new templates; v2 discovers datasets, hyperparameters, and experiments automatically.

Inside the AI Scientist v2 Agentic System

The heart of v2 is the AgentManager class, which replaces v1's static script glue. When initialized with a research idea and configuration, the manager sets up stages, journals, and sub-stage logic to drive the search process.

from ai_scientist.treesearch.agent_manager import AgentManager
from ai_scientist.utils.token_tracker import token_tracker
import json
import yaml
from pathlib import Path

# Load a research idea description (converted from markdown to JSON)

with open("ai_scientist/ideas/my_topic.json") as f:
    task_desc = json.load(f)

# Load the default configuration (bfts_config.yaml)

cfg = yaml.safe_load(open("bfts_config.yaml"))

# Create the manager – this orchestrates the multi-stage tree search

manager = AgentManager(
    task_desc=json.dumps(task_desc), 
    cfg=cfg, 
    workspace_dir=Path("experiments/run1")
)
manager.run(exec_callback=my_experiment_function)

The AgentManager.run() method implements the best-first tree search (BFTS) logic, automatically creating sub-stages and deciding when to advance to the next main stage based on experimental outcomes.

Parallel Agents and Backend Abstraction

While v1 executed steps sequentially, v2 scales through ParallelAgents (ai_scientist/treesearch/parallel_agent.py), which run pools of LLM agents concurrently for any given sub-stage. The system abstracts LLM calls through a pluggable backend architecture. You can configure OpenAI, Anthropic, or Gemini models via YAML without modifying the search logic, as implemented in files like ai_scientist/treesearch/backend/backend_anthropic.py.

Running AI Scientist v2: Practical Examples

Launching a Full Experiment

Unlike v1's multi-step template preparation, v2 executes end-to-end research with a single command via launch_scientist_bfts.py:

python launch_scientist_bfts.py \
  --load_ideas "ai_scientist/ideas/i_cant_be_later_better.json" \
  --model_writeup o1-preview-2024-09-12 \
  --model_citation gpt-4o-2024-11-20 \
  --model_review gpt-4o-2024-11-20

This script parses the idea JSON, initializes the workspace, launches the agentic tree search, aggregates plots using the VLM module, generates the manuscript via ai_scientist/perform_writeup.py, and optionally conducts automated review.

Configuring LLM Backends

The flexible backend system allows swapping models through bfts_config.yaml:

backend:
  type: "anthropic"
  model: "claude-3-5-sonnet-20240620-v1:0"
  api_key: "${AWS_ACCESS_KEY_ID}"

This abstraction enables the system to route calls to ai_scientist/treesearch/backend/backend_openai.py or custom implementations without altering the core search algorithms.

Success Rates and Trade-offs

The repository documentation warns that v2 does not necessarily produce better papers than v1. The templated approach in v1 yields higher consistency and reliability when a suitable template exists, whereas v2's broader exploratory search can generate more innovative insights but with less predictable outcomes. Success in v2 depends heavily on LLM reasoning quality and the soundness of generated ideas rather than template fidelity.

Summary

  • AI Scientist v1 uses rigid, human-written markdown templates and linear shell scripts, excelling at well-defined tasks with high consistency.
  • AI Scientist v2 replaces templates with a progressive agentic tree search orchestrated by AgentManager, enabling open-ended discovery across arbitrary ML domains.
  • Key v2 components include ParallelAgents for concurrency, a VLM for plot analysis, and a backend abstraction supporting multiple LLM providers.
  • The entry point launch_scientist_bfts.py automates the full pipeline from ideation to manuscript generation.
  • v2 trades consistency for flexibility: it generalizes without templates but produces less predictable results compared to v1's structured approach.

Frequently Asked Questions

What is the main difference between AI Scientist v1 and v2?

AI Scientist v1 relies on static templates, while v2 uses agentic search. Specifically, v1 executes predefined shell scripts that render human-authored markdown templates, whereas v2 employs an AgentManager to conduct a progressive best-first tree search that dynamically discovers experiments, datasets, and hyperparameters without template guidance.

Does AI Scientist v2 produce better research papers than v1?

Not necessarily. According to the SakanaAI/AI-Scientist-v2 documentation, v1's templated approach generates more consistent and reliable outputs when a suitable template exists. V2 prioritizes exploration and can produce innovative results, but its success depends on the LLM's reasoning capabilities and the quality of generated ideas, leading to more variable outcomes.

How does the agentic tree search in v2 work?

The search is orchestrated by the AgentManager class in ai_scientist/treesearch/agent_manager.py. It loads a research idea and configuration, then iteratively proposes implementations through ParallelAgents, evaluates experimental results, and refines the approach via a best-first tree search strategy. The system automatically manages stages, checkpoints progress, and decides when to advance based on empirical results rather than predetermined scripts.

Can I use custom LLM backends with AI Scientist v2?

Yes, through the backend abstraction layer. The system supports OpenAI, Anthropic, and Gemini models via configuration in bfts_config.yaml. You can specify provider details in files like ai_scientist/treesearch/backend/backend_anthropic.py without modifying the core search logic in agent_manager.py or parallel_agent.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →