How to Run the Full AI Scientist v2 Pipeline: A Complete End-to-End Guide

To run the full AI Scientist v2 pipeline, execute ideation with perform_ideation_temp_free.py to generate structured research ideas, then launch the complete autonomous workflow via launch_scientist_bfts.py, which orchestrates experimentation, write-up, and peer review.

The SakanaAI/AI-Scientist-v2 repository implements a fully autonomous research system that transforms high-level topics into peer-review-ready papers through a four-stage workflow. This guide covers the exact commands, configuration files, and source code entry points required to execute the complete pipeline from ideation to final manuscript generation.

Stage 1: Ideation – Generate Structured Research Ideas

The pipeline begins with idea generation using ai_scientist/perform_ideation_temp_free.py. This script consumes a Markdown "workshop" file describing your research topic and produces structured JSON ideas compatible with the downstream experimentation stage.

Required Input Format

Create a workshop file (e.g., ai_scientist/ideas/my_topic.md) containing:

  • Title: Research topic name
  • Keywords: Relevant technical terms
  • TL;DR: One-sentence summary
  • Abstract: Detailed problem description

Execution Command

python ai_scientist/perform_ideation_temp_free.py \
    --model gpt-4o-2024-05-13 \
    --workshop-file ai_scientist/ideas/my_topic.md \
    --max-num-generations 5 \
    --num-reflections 5

Key implementation details:

  • The script iteratively generates ideas using a system prompt that lists available tools including SemanticScholarSearchTool and FinalizeIdea
  • Each valid idea must contain required fields: Name, Title, Short Hypothesis, Related Work, Abstract, Experiments, and Risk Factors and Limitations
  • Set OPENAI_API_KEY and optionally S2_API_KEY for Semantic Scholar integration in your environment variables
  • Output is written to <workshop>.json in the same directory

Stage 2: Experimentation – Best-First Tree Search (BFTS)

The experimentation stage executes a best-first tree search (BFTS) over the generated ideas, spawning code agents that run experiments on GPU. This is the core compute-intensive phase of the AI Scientist v2 pipeline.

Launch Command

python launch_scientist_bfts.py \
    --load_ideas ai_scientist/ideas/my_topic.json \
    --load_code \
    --add_dataset_ref \
    --model_writeup o1-preview-2024-09-12 \
    --model_citation gpt-4o-2024-11-20 \
    --model_review gpt-4o-2024-11-20 \
    --model_agg_plots o3-mini-2025-01-31 \
    --num_cite_rounds 20

Pipeline Orchestration Details

Under the hood, launch_scientist_bfts.py performs these operations:

  1. Folder preparation: Creates a timestamped directory under experiments/ (format: date_Name_attempt_X) and converts the JSON idea to markdown via idea_to_markdown
  2. Code injection: If --load_code is specified, loads a <idea>.py file matching the JSON base name and stores contents in the idea JSON under the "Code" field
  3. Dataset reference: --add_dataset_ref concatenates hf_dataset_reference.py content to the code block for HuggingFace integration
  4. Configuration editing: edit_bfts_config_file writes a temporary config pointing the tree-search to the experiment folder (see bfts_config.yaml for default parameters)
  5. Tree search execution: perform_experiments_bfts_with_agentmanager (from ai_scientist/treesearch/perform_experiments_bfts_with_agentmanager.py) drives the agent manager that expands nodes and executes Python snippets

Critical BFTS parameters in bfts_config.yaml:

  • num_workers: Parallel agent processes
  • steps: Maximum search depth
  • max_debug_depth and debug_prob: Retry logic for failing nodes
  • num_drafts: Number of experimental drafts to generate

Stage 3: Write-Up – Assemble Results into Manuscript

After experimentation completes, the pipeline automatically aggregates results and drafts the paper using either standard or ICBINB (I Can't Believe It's Not Better) formats.

Write-Up Variants

Format Entry Point Page Limit
Standard perform_writeup in ai_scientist/perform_writeup.py 8 pages
ICBINB perform_icbinb_writeup in ai_scientist/perform_icbinb_writeup.py 4 pages

Required Models

Both functions require:

  • small_model: Cost-efficient model for drafting (e.g., gpt-4o-2024-05-13)
  • big_model: High-quality model for polishing (e.g., o1-preview-2024-09-12)
  • citations_text: Bibliography compiled by gather_citations across --num-cite-rounds iterations

The aggregate_plots function from ai_scientist/perform_plotting.py first merges experiment figures using the aggregation model (--model_agg_plots), then the write-up scripts generate a PDF (<idea>.pdf) and raw markdown manuscript in the experiment directory.

Stage 4: Review – LLM-Based Peer Assessment

Unless --skip_review is specified, the pipeline automatically evaluates the generated manuscript:

  1. PDF discovery: find_pdf_path_for_review locates the latest PDF (preferring *_final*.pdf when present)
  2. Text loading: load_paper extracts PDF content for analysis
  3. Dual review generation:

Reviews are saved as review_text.txt and review_img_cap_ref.json for human inspection.

Complete End-to-End Example

This practical sequence demonstrates the full AI Scientist v2 pipeline execution:


# 1. Create research topic description

cat > ai_scientist/ideas/efficient_transfer.md <<'EOF'
Title: Efficient Zero-Shot Transfer for Vision Transformers
Keywords: vision transformers, zero-shot, transfer learning, efficient fine-tuning
TL;DR: Propose a lightweight adapter enabling zero-shot transfer from pretrained ViT to new visual domains without backbone retraining.
Abstract:
We investigate parameter-efficient transfer methods for vision transformers...
EOF

# 2. Generate structured ideas

python ai_scientist/perform_ideation_temp_free.py \
    --model gpt-4o-2024-05-13 \
    --workshop-file ai_scientist/ideas/efficient_transfer.md \
    --max-num-generations 3 \
    --num-reflections 4

# 3. Run complete pipeline (experiments → write-up → review)

python launch_scientist_bfts.py \
    --load_ideas ai_scientist/ideas/efficient_transfer.json \
    --load_code \
    --add_dataset_ref \
    --model_writeup o1-preview-2024-09-12 \
    --model_citation gpt-4o-2024-11-20 \
    --model_review gpt-4o-2024-11-20 \
    --model_agg_plots o3-mini-2025-01-31 \
    --num_cite_rounds 15

Output Structure

Upon completion, the experiment directory contains:


experiments/2024-09-12_15-30-45_efficient_zero_shot_transfer_attempt_0/
│   idea.md                    # Markdown version of research idea

│   idea.json                  # Structured idea definition

│   writeup.pdf                # Generated manuscript

│   review_text.txt            # LLM textual review

│   review_img_cap_ref.json    # Visual element review

│   token_tracker.json         # API usage statistics

│   token_tracker_interactions.json  # Detailed interaction logs

└── logs/                      # Experiment execution logs

Token Tracking and Resource Management

Throughout the pipeline execution, ai_scientist/utils/token_tracker.py records LLM token consumption. Upon completion, save_token_tracker writes token_tracker.json and token_tracker_interactions.json to the experiment folder for cost analysis.

The pipeline includes automatic cleanup: launch_scientist_bfts.py terminates stray Python/Torch processes using psutil to prevent GPU memory leaks after tree-search completion.

Summary

  • Ideation requires perform_ideation_temp_free.py with a Markdown workshop file to generate structured JSON ideas via LLM with optional Semantic Scholar integration
  • Experimentation uses launch_scientist_bfts.py to execute best-first tree search via perform_experiments_bfts_with_agentmanager, configurable through bfts_config.yaml
  • Write-up automatically aggregates plots and drafts manuscripts using perform_writeup.py or perform_icbinb_writeup.py with separate small and large model assignments
  • Review generates LLM-based peer assessments through perform_llm_review.py and perform_vlm_review.py unless skipped with --skip_review
  • Environment setup requires OPENAI_API_KEY (or model-specific keys) and optional S2_API_KEY, with GPU memory monitored via num_workers in BFTS configuration

Frequently Asked Questions

What API keys are required to run the AI Scientist v2 pipeline?

You must set OPENAI_API_KEY for OpenAI models, GEMINI_API_KEY for Gemini models, and HUGGINGFACE_API_KEY when using DeepCoder. For literature search functionality during ideation, configure S2_API_KEY for Semantic Scholar access. The pipeline validates these environment variables before initiating LLM calls in ai_scientist/llm.py.

How do I prevent GPU out-of-memory errors during experimentation?

Reduce parallel agent execution by lowering num_workers in bfts_config.yaml, or decrease steps to limit search depth. The max_debug_depth and debug_prob parameters control retry aggressiveness for failed nodes—reducing these prevents excessive concurrent GPU processes. The script automatically cleans up stray Python/Torch processes post-execution using psutil to free VRAM.

Can I skip specific stages of the pipeline?

Yes. Use --skip_writeup to halt after experimentation, or --skip_review to omit the peer-review stage. These flags are processed in launch_scientist_bfts.py, allowing you to debug the BFTS phase independently before generating manuscripts. Each run creates a fresh timestamped folder, so rerunning with different flags never overwrites previous results.

How does the pipeline handle code injection for specific research ideas?

When --load_code is specified, launch_scientist_bfts.py searches for a Python file matching the idea JSON basename (e.g., my_idea.py for my_idea.json) and injects its contents into the idea JSON under the "Code" field. The --add_dataset_ref flag additionally concatenates hf_dataset_reference.py for HuggingFace dataset integration, enabling reproducible experiment templates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →