How to Run the Full AI Scientist v2 Pipeline: A Complete End-to-End Guide
To run the full AI Scientist v2 pipeline, execute ideation with perform_ideation_temp_free.py to generate structured research ideas, then launch the complete autonomous workflow via launch_scientist_bfts.py, which orchestrates experimentation, write-up, and peer review.
The SakanaAI/AI-Scientist-v2 repository implements a fully autonomous research system that transforms high-level topics into peer-review-ready papers through a four-stage workflow. This guide covers the exact commands, configuration files, and source code entry points required to execute the complete pipeline from ideation to final manuscript generation.
Stage 1: Ideation – Generate Structured Research Ideas
The pipeline begins with idea generation using ai_scientist/perform_ideation_temp_free.py. This script consumes a Markdown "workshop" file describing your research topic and produces structured JSON ideas compatible with the downstream experimentation stage.
Required Input Format
Create a workshop file (e.g., ai_scientist/ideas/my_topic.md) containing:
- Title: Research topic name
- Keywords: Relevant technical terms
- TL;DR: One-sentence summary
- Abstract: Detailed problem description
Execution Command
python ai_scientist/perform_ideation_temp_free.py \
--model gpt-4o-2024-05-13 \
--workshop-file ai_scientist/ideas/my_topic.md \
--max-num-generations 5 \
--num-reflections 5
Key implementation details:
- The script iteratively generates ideas using a system prompt that lists available tools including
SemanticScholarSearchToolandFinalizeIdea - Each valid idea must contain required fields:
Name,Title,Short Hypothesis,Related Work,Abstract,Experiments, andRisk Factors and Limitations - Set
OPENAI_API_KEYand optionallyS2_API_KEYfor Semantic Scholar integration in your environment variables - Output is written to
<workshop>.jsonin the same directory
Stage 2: Experimentation – Best-First Tree Search (BFTS)
The experimentation stage executes a best-first tree search (BFTS) over the generated ideas, spawning code agents that run experiments on GPU. This is the core compute-intensive phase of the AI Scientist v2 pipeline.
Launch Command
python launch_scientist_bfts.py \
--load_ideas ai_scientist/ideas/my_topic.json \
--load_code \
--add_dataset_ref \
--model_writeup o1-preview-2024-09-12 \
--model_citation gpt-4o-2024-11-20 \
--model_review gpt-4o-2024-11-20 \
--model_agg_plots o3-mini-2025-01-31 \
--num_cite_rounds 20
Pipeline Orchestration Details
Under the hood, launch_scientist_bfts.py performs these operations:
- Folder preparation: Creates a timestamped directory under
experiments/(format:date_Name_attempt_X) and converts the JSON idea to markdown viaidea_to_markdown - Code injection: If
--load_codeis specified, loads a<idea>.pyfile matching the JSON base name and stores contents in the idea JSON under the"Code"field - Dataset reference:
--add_dataset_refconcatenateshf_dataset_reference.pycontent to the code block for HuggingFace integration - Configuration editing:
edit_bfts_config_filewrites a temporary config pointing the tree-search to the experiment folder (seebfts_config.yamlfor default parameters) - Tree search execution:
perform_experiments_bfts_with_agentmanager(fromai_scientist/treesearch/perform_experiments_bfts_with_agentmanager.py) drives the agent manager that expands nodes and executes Python snippets
Critical BFTS parameters in bfts_config.yaml:
num_workers: Parallel agent processessteps: Maximum search depthmax_debug_depthanddebug_prob: Retry logic for failing nodesnum_drafts: Number of experimental drafts to generate
Stage 3: Write-Up – Assemble Results into Manuscript
After experimentation completes, the pipeline automatically aggregates results and drafts the paper using either standard or ICBINB (I Can't Believe It's Not Better) formats.
Write-Up Variants
| Format | Entry Point | Page Limit |
|---|---|---|
| Standard | perform_writeup in ai_scientist/perform_writeup.py |
8 pages |
| ICBINB | perform_icbinb_writeup in ai_scientist/perform_icbinb_writeup.py |
4 pages |
Required Models
Both functions require:
small_model: Cost-efficient model for drafting (e.g.,gpt-4o-2024-05-13)big_model: High-quality model for polishing (e.g.,o1-preview-2024-09-12)citations_text: Bibliography compiled bygather_citationsacross--num-cite-roundsiterations
The aggregate_plots function from ai_scientist/perform_plotting.py first merges experiment figures using the aggregation model (--model_agg_plots), then the write-up scripts generate a PDF (<idea>.pdf) and raw markdown manuscript in the experiment directory.
Stage 4: Review – LLM-Based Peer Assessment
Unless --skip_review is specified, the pipeline automatically evaluates the generated manuscript:
- PDF discovery:
find_pdf_path_for_reviewlocates the latest PDF (preferring*_final*.pdfwhen present) - Text loading:
load_paperextracts PDF content for analysis - Dual review generation:
perform_review(fromai_scientist/perform_llm_review.py) conducts textual analysisperform_imgs_cap_ref_review(fromai_scientist/perform_vlm_review.py) reviews figure captions and references
Reviews are saved as review_text.txt and review_img_cap_ref.json for human inspection.
Complete End-to-End Example
This practical sequence demonstrates the full AI Scientist v2 pipeline execution:
# 1. Create research topic description
cat > ai_scientist/ideas/efficient_transfer.md <<'EOF'
Title: Efficient Zero-Shot Transfer for Vision Transformers
Keywords: vision transformers, zero-shot, transfer learning, efficient fine-tuning
TL;DR: Propose a lightweight adapter enabling zero-shot transfer from pretrained ViT to new visual domains without backbone retraining.
Abstract:
We investigate parameter-efficient transfer methods for vision transformers...
EOF
# 2. Generate structured ideas
python ai_scientist/perform_ideation_temp_free.py \
--model gpt-4o-2024-05-13 \
--workshop-file ai_scientist/ideas/efficient_transfer.md \
--max-num-generations 3 \
--num-reflections 4
# 3. Run complete pipeline (experiments → write-up → review)
python launch_scientist_bfts.py \
--load_ideas ai_scientist/ideas/efficient_transfer.json \
--load_code \
--add_dataset_ref \
--model_writeup o1-preview-2024-09-12 \
--model_citation gpt-4o-2024-11-20 \
--model_review gpt-4o-2024-11-20 \
--model_agg_plots o3-mini-2025-01-31 \
--num_cite_rounds 15
Output Structure
Upon completion, the experiment directory contains:
experiments/2024-09-12_15-30-45_efficient_zero_shot_transfer_attempt_0/
│ idea.md # Markdown version of research idea
│ idea.json # Structured idea definition
│ writeup.pdf # Generated manuscript
│ review_text.txt # LLM textual review
│ review_img_cap_ref.json # Visual element review
│ token_tracker.json # API usage statistics
│ token_tracker_interactions.json # Detailed interaction logs
└── logs/ # Experiment execution logs
Token Tracking and Resource Management
Throughout the pipeline execution, ai_scientist/utils/token_tracker.py records LLM token consumption. Upon completion, save_token_tracker writes token_tracker.json and token_tracker_interactions.json to the experiment folder for cost analysis.
The pipeline includes automatic cleanup: launch_scientist_bfts.py terminates stray Python/Torch processes using psutil to prevent GPU memory leaks after tree-search completion.
Summary
- Ideation requires
perform_ideation_temp_free.pywith a Markdown workshop file to generate structured JSON ideas via LLM with optional Semantic Scholar integration - Experimentation uses
launch_scientist_bfts.pyto execute best-first tree search viaperform_experiments_bfts_with_agentmanager, configurable throughbfts_config.yaml - Write-up automatically aggregates plots and drafts manuscripts using
perform_writeup.pyorperform_icbinb_writeup.pywith separate small and large model assignments - Review generates LLM-based peer assessments through
perform_llm_review.pyandperform_vlm_review.pyunless skipped with--skip_review - Environment setup requires
OPENAI_API_KEY(or model-specific keys) and optionalS2_API_KEY, with GPU memory monitored vianum_workersin BFTS configuration
Frequently Asked Questions
What API keys are required to run the AI Scientist v2 pipeline?
You must set OPENAI_API_KEY for OpenAI models, GEMINI_API_KEY for Gemini models, and HUGGINGFACE_API_KEY when using DeepCoder. For literature search functionality during ideation, configure S2_API_KEY for Semantic Scholar access. The pipeline validates these environment variables before initiating LLM calls in ai_scientist/llm.py.
How do I prevent GPU out-of-memory errors during experimentation?
Reduce parallel agent execution by lowering num_workers in bfts_config.yaml, or decrease steps to limit search depth. The max_debug_depth and debug_prob parameters control retry aggressiveness for failed nodes—reducing these prevents excessive concurrent GPU processes. The script automatically cleans up stray Python/Torch processes post-execution using psutil to free VRAM.
Can I skip specific stages of the pipeline?
Yes. Use --skip_writeup to halt after experimentation, or --skip_review to omit the peer-review stage. These flags are processed in launch_scientist_bfts.py, allowing you to debug the BFTS phase independently before generating manuscripts. Each run creates a fresh timestamped folder, so rerunning with different flags never overwrites previous results.
How does the pipeline handle code injection for specific research ideas?
When --load_code is specified, launch_scientist_bfts.py searches for a Python file matching the idea JSON basename (e.g., my_idea.py for my_idea.json) and injects its contents into the idea JSON under the "Code" field. The --add_dataset_ref flag additionally concatenates hf_dataset_reference.py for HuggingFace dataset integration, enabling reproducible experiment templates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →