What Are the Dependencies for AI Scientist v2? Complete Package Guide
The AI Scientist v2 system from SakanaAI requires a curated ecosystem of 30+ Python packages spanning LLM APIs, machine learning frameworks, visualization tools, and cloud services, all centralized in requirements.txt for one-command installation.
SakanaAI/AI-Scientist-v2 is an open-source autonomous research framework that generates scientific papers using large language models. The dependencies for AI Scientist v2 enable everything from API communication with Anthropic and OpenAI to distributed storage on AWS S3, forming the technical backbone of the automated experimentation pipeline.
LLM API and Client Libraries
The system interacts with frontier language models through dedicated client libraries wrapped with robust retry mechanisms.
openaiandanthropic: Provide native SDK access to GPT-4, Claude, and other frontier models. These are imported and configured inai_scientist/llm.pyto instantiate chat clients.backoff: Implements exponential backoff logic around API calls to handle rate limits and transient network failures gracefully.tiktoken: Counts tokens in prompts and responses for usage tracking and cost estimation before sending requests to OpenAI endpoints.
Machine Learning and Data Processing
Core numerical and ML operations rely on industry-standard libraries distributed across the ai_scientist/treesearch/ modules and experiment runners.
numpy: Powers numerical scoring and array operations within tree search algorithms.transformersanddatasets: Enable on-the-fly loading and fine-tuning of pre-trained Hugging Face models during experimentation.wandb: Integrates Weights & Biases for experiment tracking and metric logging.tqdmandrich: Render live progress bars and styled console output during long-running research loops.humanize: Formats timestamps and byte sizes into human-readable strings for logging.
Visualization and Document Generation
Publication-ready figures and PDF handling require specialized plotting and document parsing tools utilized in ai_scientist/perform_plotting and ai_scientist/perform_writeup.
matplotlibandseaborn: Generate line plots, histograms, and statistical visualizations from experiment DataFrames.pypdfandpymupdf4llm: Parse existing research papers and extract text from PDFs for context injection into LLM prompts.
Cloud Infrastructure and Storage
Distributed experiments persist artifacts to object storage using AWS SDKs referenced throughout ai_scientist/utils/.
boto3andbotocore: Interface with AWS S3 to upload model checkpoints, JSON results, and generated LaTeX papers to remote buckets.
Configuration and Utility Tooling
Supporting infrastructure for code quality, graph processing, and typed configurations appears across the codebase.
omegaconf: Manages hierarchical YAML configurations with type safety.dataclasses-json: Serializes experiment state objects to JSON for reproducibility.python-igraph: Implements graph algorithms for the tree-search planning logic inai_scientist/treesearch/.black,jsonschema,genson,funcy,coolname, andshutup: Handle code formatting, schema validation, functional programming utilities, random name generation, and suppression of superfluous warnings.
Architecture Integration
In ai_scientist/llm.py, the create_client function combines openai/anthropic clients with backoff retry wrappers and tiktoken accounting. The tree search orchestration in ai_scientist/treesearch/ leverages numpy for numerical operations, python-igraph for maintaining search trees, and rich/tqdm for real-time status displays. Cloud persistence layers in ai_scientist/utils/ use boto3 to stream large artifacts to S3 without blocking the main research loop.
Installation and Setup
Install all dependencies from the repository root using the centralized requirements file:
pip install -r requirements.txt
This command installs the complete dependency graph including transitive requirements for PyTorch, Hugging Face ecosystems, and AWS tooling.
Practical Implementation Examples
Initializing an LLM Client with Retry Logic
The create_client wrapper in ai_scientist/llm.py abstracts raw SDK initialization and automatically configures backoff retries:
from ai_scientist.llm import create_client
client = create_client(provider="openai", model="gpt-4o-mini")
response = client.chat.complete(
messages=[{"role": "user", "content": "Explain the significance of the Fourier transform."}]
)
print(response["content"])
Generating Research Visualizations
Experiment metrics convert to publication figures using the visualization stack:
import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd
# Assume df contains experiment metrics
sns.lineplot(data=df, x="epoch", y="accuracy", hue="run_id")
plt.title("Training Accuracy Over Epochs")
plt.savefig("results/accuracy_plot.png")
Persisting Artifacts to AWS S3
The cloud storage integration enables durable backup of research outputs:
import boto3
import json
s3 = boto3.client("s3")
bucket = "ai-scientist-artifacts"
payload = json.dumps({"run_id": "abc123", "metrics": {"accuracy": 0.92}})
s3.put_object(Bucket=bucket, Key="runs/abc123/results.json", Body=payload)
Summary
- AI Scientist v2 dependencies are declared in
requirements.txtand installable viapip install -r requirements.txt. - LLM interaction relies on
openai,anthropic,backoff, andtiktokenfor robust API communication with retry logic. - Machine learning workflows use
numpy,transformers,datasets, andwandbfor model handling and experiment tracking. - Visualization and document processing depend on
matplotlib,seaborn,pypdf, andpymupdf4llmfor figure generation and PDF parsing. - Cloud persistence requires
boto3andbotocorefor S3 storage of research artifacts. - Configuration and search algorithms utilize
omegaconf,dataclasses-json, andpython-igraphfor typed configs and tree-based planning.
Frequently Asked Questions
How do I install all dependencies for AI Scientist v2?
Run pip install -r requirements.txt from the repository root directory. This installs the complete Python ecosystem including LLM clients, ML frameworks, and AWS SDKs required by launch_scientist_bfts.py and supporting modules.
Which packages handle LLM API retries and token counting?
The backoff library implements exponential retry logic for flaky API calls, while tiktoken performs pre-flight token counting. Both wrap the openai and anthropic clients inside ai_scientist/llm.py to ensure reliable communication with rate-limited endpoints.
What visualization libraries does AI Scientist v2 use for research figures?
The system uses matplotlib and seaborn for statistical plotting and pypdf/pymupdf4llm for PDF text extraction. These appear in the writeup and plotting modules to generate publication-quality figures and parse existing literature.
How does AI Scientist v2 manage cloud storage of experiment artifacts?
The boto3 and botocore packages provide AWS S3 integration within ai_scientist/utils/, enabling the system to serialize JSON results and upload large model checkpoints to remote storage buckets during distributed experiments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →