How to Set Up the Environment for AI Scientist v2: Complete Installation Guide

To set up the environment for AI Scientist v2, create a Python 3.11 Conda environment, install CUDA-enabled PyTorch, add PDF/LaTeX utilities, install requirements.txt dependencies, and configure API keys for your chosen LLM providers.

The SakanaAI/AI-Scientist-v2 repository is a Python-based research automation framework that leverages large language models (LLMs) to generate scientific hypotheses, execute experiments, and produce academic manuscripts. Setting up this environment requires specific GPU-accelerated libraries, scientific writing tools for PDF generation, and authentication tokens for multiple LLM backends. Follow the steps below to configure your local machine for running the full ideation and experiment pipelines.

Step 1: Create a Fresh Conda Environment

AI Scientist v2 requires Python 3.11 for compatibility with its dependency tree. Create an isolated Conda environment to avoid conflicts with existing packages.

conda create -n ai_scientist python=3.11
conda activate ai_scientist

These commands establish the ai_scientist environment, as specified in the repository's README (lines 45-48).

Step 2: Install PyTorch with CUDA Support

The framework relies on GPU-accelerated PyTorch for model inference and data processing. Install PyTorch with CUDA support, adjusting the CUDA version (12.4 in the example below) to match your NVIDIA driver.

conda install pytorch torchvision torchaudio pytorch-cuda=12.4 -c pytorch -c nvidia

This installation targets CUDA 12.4; modify pytorch-cuda=12.4 if your system requires a different version (see README lines 50-52).

Step 3: Add PDF and LaTeX Utilities

Manuscript generation requires system-level tools for processing PDFs and validating LaTeX syntax. Install Poppler for PDF manipulation and ChkTeX for LaTeX linting.

conda install anaconda::poppler
conda install conda-forge::chktex

These utilities enable the automated paper-writing pipeline to render figures and compile academic documents (README lines 53-56).

Step 4: Install Python Dependencies

With the environment and system tools ready, install the remaining Python packages specified in the repository's requirements file.

pip install -r requirements.txt

This command installs the full dependency tree including web search utilities, data processing libraries, and LLM client interfaces (README line 58).

Step 5: Configure LLM API Keys

AI Scientist v2 supports multiple LLM backends. Export the relevant API keys for the providers you intend to use:

  • OpenAI models: export OPENAI_API_KEY="YOUR_OPENAI_KEY"
  • Google Gemini: export GEMINI_API_KEY="YOUR_GEMINI_KEY"
  • Claude via AWS Bedrock: export AWS_ACCESS_KEY_ID="...", export AWS_SECRET_ACCESS_KEY="...", export AWS_REGION_NAME="..."
  • Semantic Scholar (optional, for higher throughput): export S2_API_KEY="YOUR_SEMANTIC_SCHOLAR_KEY"

If using Claude through AWS Bedrock, install the supplemental client:

pip install "anthropic[bedrock]"

API configuration details are documented in README lines 66-84. The framework uses these credentials in ai_scientist/treesearch/backend/backend_openai.py for OpenAI-compatible calls and ai_scientist/tools/semantic_scholar.py for literature searches.

Step 6: Verify Installation with a Test Run

Confirm your setup by running the ideation pipeline, which generates research ideas without consuming significant GPU resources.

python ai_scientist/perform_ideation_temp_free.py \
    --workshop-file "ai_scientist/ideas/my_research_topic.md" \
    --model gpt-4o-2024-05-13 \
    --max-num-generations 20 \
    --num-reflections 5

This script creates a JSON file containing structured research ideas (README lines 99-112).

To execute a full experiment including tree-search and manuscript generation, run:

python launch_scientist_bfts.py \
    --load_ideas "ai_scientist/ideas/my_research_topic.json" \
    --load_code \
    --add_dataset_ref \
    --model_writeup o1-preview-2024-09-12 \
    --model_citation gpt-4o-2024-11-20 \
    --model_review gpt-4o-2024-11-20 \
    --model_agg_plots o3-mini-2025-01-31 \
    --num_cite_rounds 20

This entry point orchestrates the agentic best-first tree search, experimental validation, and PDF compilation (README lines 143-152).

Key Files in the Repository

Understanding these core files helps troubleshoot setup issues:

Summary

Setting up the environment for AI Scientist v2 involves six critical stages:

  • Create a dedicated Python 3.11 Conda environment named ai_scientist.
  • Install CUDA-enabled PyTorch matching your GPU driver version.
  • Add system-level PDF and LaTeX tools (poppler, chktex) for manuscript generation.
  • Install all Python dependencies via requirements.txt.
  • Configure API keys for OpenAI, Gemini, AWS Bedrock (optional), and Semantic Scholar (optional).
  • Verify the installation by running perform_ideation_temp_free.py before executing full experiments.

Frequently Asked Questions

Do I need a GPU to run AI Scientist v2?

Yes. The framework requires CUDA-enabled PyTorch for model inference and experimental computation. While some ideation steps may run on CPU, the full pipeline—including the best-first tree search in launch_scientist_bfts.py—expects GPU acceleration for acceptable performance.

Can I use only OpenAI models and skip AWS configuration?

Yes. AWS credentials are only required if you intend to use Claude models via AWS Bedrock. If you use OpenAI (gpt-4o, o1-preview, etc.) or Google Gemini models exclusively, you only need to set OPENAI_API_KEY or GEMINI_API_KEY. The code automatically detects available backends based on which environment variables are present.

What if my CUDA version differs from 12.4?

Adjust the PyTorch installation command to match your system's CUDA version. Replace pytorch-cuda=12.4 with your installed version (e.g., 11.8 or 12.1). PyTorch maintains compatibility matrices on their official site; verify your combination of Python, PyTorch, and CUDA before installing.

Where are the generated research papers saved?

After a successful run of launch_scientist_bfts.py, the system saves PDF manuscripts and LaTeX source files in the output directory specified by the experiment configuration (defaulting to a timestamped folder under the project root). The bfts_config.yaml file controls these path behaviors and naming conventions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →