How to Generate Research Papers with AI Scientist v2: End-to-End Workflow Guide

AI Scientist v2 is an open-source autonomous research system that generates full scientific manuscripts through four automated phases—ideation, agentic tree-search experimentation, citation gathering, and structured write-up—controlled by the single entry point launch_scientist_bfts.py.

AI Scientist v2 from SakanaAI is a modular, end-to-end system for autonomously generating research papers. To generate research papers with AI Scientist v2, you execute a two-stage pipeline that moves from topic description to peer-review-ready PDF without manual intervention. The codebase implements a Best-First Tree Search (BFTS) agentic architecture that explores experimental ideas, validates them through code execution, and compiles results into formatted manuscripts.

The Four-Phase Automated Research Pipeline

The system operates through four distinct phases orchestrated by launch_scientist_bfts.py:

The pipeline begins in ai_scientist/perform_ideation_temp_free.py. This module reads a workshop file—a Markdown document containing Title, Keywords, TL;DR, and Abstract—that defines your research domain.

The script invokes an LLM via ai_scientist.llm.create_client with a system prompt enumerating available tools. The primary tool is SemanticScholarSearchTool (implemented in ai_scientist/tools/semantic_scholar.py), which the model calls iteratively during a reflection loop to verify novelty against prior work.

After reflection, the system emits a JSON-structured IDEA object containing Name, Title, Hypothesis, Related Work, Abstract, Experiments, and Risk Factors. The output persists to <workshop>.json for the next stage.

The experimentation phase uses perform_experiments_bfts_with_agentmanager.py to implement BFTS. The entry point launch_scientist_bfts.py converts the selected idea JSON into a Markdown file via idea_to_markdown (from bfts_utils.py) and updates bfts_config.yaml to point to a unique output folder using edit_bfts_config_file.

The Agent Manager spawns a pool of worker agents (controlled by num_workers) that explore the experiment tree. Each node consists of three operations:

  1. Code Generation: An LLM writes experiment code based on the current hypothesis.
  2. Execution: The generated script runs inside a sandbox environment.
  3. Result Evaluation: A model scores the outcome and decides whether to expand the node or prune the branch.

The search proceeds until the steps budget exhausts or the experiment succeeds. Results accumulate in experiment_results/ with an interactive unified_tree_viz.html visualization.

Phase 3: Citation Gathering and Manuscript Drafting

After experimentation, the system transitions to writing. The citation phase runs a multi-round search (specified by --num_cite_rounds) using the LLM configured via --model_citation. Each round can invoke the Semantic Scholar tool to fetch DOI, title, and abstract data, concatenated into citations_text.

The drafting process employs a two-model pipeline:

  • Small Model (--model_writeup_small): Creates a first-pass draft including outline and figures.
  • Large Model (--model_writeup): Rewrites the draft into a polished PDF respecting page limits.

For short-form output (4 pages), perform_icbinb_writeup.py handles the condensed format, while perform_writeup.py generates the standard 8-page manuscript. The system retries the write-up step up to --writeup-retries times to ensure robustness.

Phase 4: Automated Peer Review and Validation

The final phase simulates academic peer review. perform_llm_review.py (text review) and perform_vlm_review.py (vision-language review) process the generated PDF using the model specified by --model_review.

The VLM component (perform_imgs_cap_ref_review) specifically extracts figure captions and suggests improvements, outputting review_img_cap_ref.json. Meanwhile, token_tracker.py records all LLM token usage throughout the pipeline, saving summaries as token_tracker.json and token_tracker_interactions.json. Upon completion, a cleanup routine terminates stray PyTorch and Multiprocessing processes to prevent GPU memory leaks.

Executing the Complete Pipeline

To generate research papers with AI Scientist v2, you typically run two commands: one for ideation and one for the full experiment-to-paper pipeline.

Step 1: Generate Research Ideas

Create a Markdown workshop file describing your research topic (see ai_scientist/ideas/i_cant_believe_its_not_better.md for an example), then run:

python ai_scientist/perform_ideation_temp_free.py \
    --model gpt-4o-2024-05-13 \
    --max-num-generations 5 \
    --num-reflections 5 \
    --workshop-file ai_scientist/ideas/my_topic.md

This produces my_topic.json containing up to 5 structured research ideas.

Step 2: Run Full Experimentation and Write-up

Execute the complete pipeline using launch_scientist_bfts.py:

python launch_scientist_bfts.py \
    --load_ideas ai_scientist/ideas/my_topic.json \
    --load_code \
    --add_dataset_ref \
    --model_writeup o1-preview-2024-09-12 \
    --model_citation gpt-4o-2024-11-20 \
    --model_review gpt-4o-2024-11-20 \
    --model_agg_plots o3-mini-2025-01-31 \
    --num_cite_rounds 20 \
    --writeup-retries 3

Key flags include --load_code (to merge accompanying Python files), --add_dataset_ref (to inject HuggingFace dataset stubs), and GPU allocation via --gpu-ids.

Output Directory Structure

Results are stored in experiments/<timestamp>_<idea_name>_attempt_0/:

  • idea.md: Markdown representation of the selected research idea.
  • experiment_results/: Raw logs and artifacts from each BFTS node.
  • unified_tree_viz.html: Interactive visualization of the experiment tree.
  • paper.pdf or paper_4page.pdf: The final formatted manuscript.
  • review_text.txt and review_img_cap_ref.json: Automated review feedback.
  • token_tracker*.json: Complete token usage accounting.

Configuration and Key Source Files

The following files constitute the core architecture for generating research papers with AI Scientist v2:

File Role
launch_scientist_bfts.py Top-level orchestrator that wires configuration, prepares directories, and executes the full pipeline.
perform_ideation_temp_free.py Template-free idea generation with reflection loops and Semantic Scholar integration.
perform_experiments_bfts_with_agentmanager.py Implements BFTS with parallel agent workers for code generation and evaluation.
bfts_config.yaml Default configuration for tree search parameters including worker count and step limits.
perform_writeup.py Generates 8-page manuscripts using the small-then-large model pipeline.
perform_icbinb_writeup.py Generates condensed 4-page workshop papers with gather_citations integration.
perform_llm_review.py Produces textual peer reviews of the final manuscript.
perform_vlm_review.py Vision-language review for figure captions and image quality.
semantic_scholar.py API wrapper for literature search and citation metadata retrieval.
token_tracker.py Centralized bookkeeping of LLM token consumption across all phases.
bfts_utils.py Utility functions including idea_to_markdown and edit_bfts_config_file.

Summary

  • AI Scientist v2 automates the full research lifecycle through four phases: ideation, BFTS experimentation, citation-enhanced writing, and simulated peer review.
  • The entry point launch_scientist_bfts.py coordinates all components, from loading ideas to final PDF generation.
  • Best-First Tree Search in perform_experiments_bfts_with_agentmanager.py enables parallel, agentic exploration of experimental hypotheses with automatic pruning.
  • Dual-model drafting uses a small model for outlines and a large model (e.g., o1-preview) for final manuscript polish, with retry logic for robustness.
  • Complete provenance is maintained through token_tracker.py and automated review generation, ensuring transparency in the autonomous research process.

Frequently Asked Questions

What models work best with AI Scientist v2 for paper generation?

OpenAI's o1-preview and gpt-4o variants are currently the primary supported models for write-up and citation tasks, as implemented in launch_scientist_bfts.py. The system uses o1-preview-2024-09-12 or similar for high-quality manuscript generation (--model_writeup), while gpt-4o-2024-11-20 handles citation gathering (--model_citation) and review tasks (--model_review). You can specify different models for plotting aggregation via --model_agg_plots (e.g., o3-mini-2025-01-31).

How does the Best-First Tree Search (BFTS) avoid wasting compute on bad ideas?

The BFTS implementation in perform_experiments_bfts_with_agentmanager.py prunes failing branches through continuous evaluation. Each tree node undergoes code generation, execution, and result scoring. The evaluation model decides whether to expand a node further or terminate that branch. This agentic approach, configured via bfts_config.yaml, ensures only promising experimental paths consume your GPU budget defined by the steps parameter.

Can I customize the output format for different venues?

Yes, AI Scientist v2 supports both standard 8-page and condensed 4-page formats. The perform_writeup.py module generates full-length manuscripts, while perform_icbinb_writeup.py produces the 4-page "I Can't Believe It's Not Better" workshop format. Both use the same underlying citation gathering and dual-model drafting pipeline, differing only in length constraints and section compression.

Where is token usage tracked for cost monitoring?

All token consumption is recorded by token_tracker.py and saved as JSON in the experiment directory. The system generates token_tracker.json (aggregated totals) and token_tracker_interactions.json (detailed per-call logs) automatically. This integration runs throughout ideation, experimentation, write-up, and review phases, allowing precise cost accounting for each generated paper.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →