What Are the Dependencies for AI Scientist v2? Complete Package Guide

The AI Scientist v2 system from SakanaAI requires a curated ecosystem of 30+ Python packages spanning LLM APIs, machine learning frameworks, visualization tools, and cloud services, all centralized in requirements.txt for one-command installation.

SakanaAI/AI-Scientist-v2 is an open-source autonomous research framework that generates scientific papers using large language models. The dependencies for AI Scientist v2 enable everything from API communication with Anthropic and OpenAI to distributed storage on AWS S3, forming the technical backbone of the automated experimentation pipeline.

LLM API and Client Libraries

The system interacts with frontier language models through dedicated client libraries wrapped with robust retry mechanisms.

  • openai and anthropic: Provide native SDK access to GPT-4, Claude, and other frontier models. These are imported and configured in ai_scientist/llm.py to instantiate chat clients.
  • backoff: Implements exponential backoff logic around API calls to handle rate limits and transient network failures gracefully.
  • tiktoken: Counts tokens in prompts and responses for usage tracking and cost estimation before sending requests to OpenAI endpoints.

Machine Learning and Data Processing

Core numerical and ML operations rely on industry-standard libraries distributed across the ai_scientist/treesearch/ modules and experiment runners.

  • numpy: Powers numerical scoring and array operations within tree search algorithms.
  • transformers and datasets: Enable on-the-fly loading and fine-tuning of pre-trained Hugging Face models during experimentation.
  • wandb: Integrates Weights & Biases for experiment tracking and metric logging.
  • tqdm and rich: Render live progress bars and styled console output during long-running research loops.
  • humanize: Formats timestamps and byte sizes into human-readable strings for logging.

Visualization and Document Generation

Publication-ready figures and PDF handling require specialized plotting and document parsing tools utilized in ai_scientist/perform_plotting and ai_scientist/perform_writeup.

  • matplotlib and seaborn: Generate line plots, histograms, and statistical visualizations from experiment DataFrames.
  • pypdf and pymupdf4llm: Parse existing research papers and extract text from PDFs for context injection into LLM prompts.

Cloud Infrastructure and Storage

Distributed experiments persist artifacts to object storage using AWS SDKs referenced throughout ai_scientist/utils/.

  • boto3 and botocore: Interface with AWS S3 to upload model checkpoints, JSON results, and generated LaTeX papers to remote buckets.

Configuration and Utility Tooling

Supporting infrastructure for code quality, graph processing, and typed configurations appears across the codebase.

  • omegaconf: Manages hierarchical YAML configurations with type safety.
  • dataclasses-json: Serializes experiment state objects to JSON for reproducibility.
  • python-igraph: Implements graph algorithms for the tree-search planning logic in ai_scientist/treesearch/.
  • black, jsonschema, genson, funcy, coolname, and shutup: Handle code formatting, schema validation, functional programming utilities, random name generation, and suppression of superfluous warnings.

Architecture Integration

In ai_scientist/llm.py, the create_client function combines openai/anthropic clients with backoff retry wrappers and tiktoken accounting. The tree search orchestration in ai_scientist/treesearch/ leverages numpy for numerical operations, python-igraph for maintaining search trees, and rich/tqdm for real-time status displays. Cloud persistence layers in ai_scientist/utils/ use boto3 to stream large artifacts to S3 without blocking the main research loop.

Installation and Setup

Install all dependencies from the repository root using the centralized requirements file:

pip install -r requirements.txt

This command installs the complete dependency graph including transitive requirements for PyTorch, Hugging Face ecosystems, and AWS tooling.

Practical Implementation Examples

Initializing an LLM Client with Retry Logic

The create_client wrapper in ai_scientist/llm.py abstracts raw SDK initialization and automatically configures backoff retries:

from ai_scientist.llm import create_client

client = create_client(provider="openai", model="gpt-4o-mini")

response = client.chat.complete(
    messages=[{"role": "user", "content": "Explain the significance of the Fourier transform."}]
)
print(response["content"])

Generating Research Visualizations

Experiment metrics convert to publication figures using the visualization stack:

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

# Assume df contains experiment metrics

sns.lineplot(data=df, x="epoch", y="accuracy", hue="run_id")
plt.title("Training Accuracy Over Epochs")
plt.savefig("results/accuracy_plot.png")

Persisting Artifacts to AWS S3

The cloud storage integration enables durable backup of research outputs:

import boto3
import json

s3 = boto3.client("s3")
bucket = "ai-scientist-artifacts"

payload = json.dumps({"run_id": "abc123", "metrics": {"accuracy": 0.92}})
s3.put_object(Bucket=bucket, Key="runs/abc123/results.json", Body=payload)

Summary

  • AI Scientist v2 dependencies are declared in requirements.txt and installable via pip install -r requirements.txt.
  • LLM interaction relies on openai, anthropic, backoff, and tiktoken for robust API communication with retry logic.
  • Machine learning workflows use numpy, transformers, datasets, and wandb for model handling and experiment tracking.
  • Visualization and document processing depend on matplotlib, seaborn, pypdf, and pymupdf4llm for figure generation and PDF parsing.
  • Cloud persistence requires boto3 and botocore for S3 storage of research artifacts.
  • Configuration and search algorithms utilize omegaconf, dataclasses-json, and python-igraph for typed configs and tree-based planning.

Frequently Asked Questions

How do I install all dependencies for AI Scientist v2?

Run pip install -r requirements.txt from the repository root directory. This installs the complete Python ecosystem including LLM clients, ML frameworks, and AWS SDKs required by launch_scientist_bfts.py and supporting modules.

Which packages handle LLM API retries and token counting?

The backoff library implements exponential retry logic for flaky API calls, while tiktoken performs pre-flight token counting. Both wrap the openai and anthropic clients inside ai_scientist/llm.py to ensure reliable communication with rate-limited endpoints.

What visualization libraries does AI Scientist v2 use for research figures?

The system uses matplotlib and seaborn for statistical plotting and pypdf/pymupdf4llm for PDF text extraction. These appear in the writeup and plotting modules to generate publication-quality figures and parse existing literature.

How does AI Scientist v2 manage cloud storage of experiment artifacts?

The boto3 and botocore packages provide AWS S3 integration within ai_scientist/utils/, enabling the system to serialize JSON results and upload large model checkpoints to remote storage buckets during distributed experiments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →