# What Are the Dependencies for AI Scientist v2? Complete Package Guide

> Discover the dependencies for AI Scientist v2. This guide details the 30+ Python packages needed, including LLM APIs and ML frameworks, all installable with requirements.txt.

- Repository: [Sakana AI/AI-Scientist-v2](https://github.com/SakanaAI/AI-Scientist-v2)
- Tags: getting-started
- Published: 2026-03-28

---

**The AI Scientist v2 system from SakanaAI requires a curated ecosystem of 30+ Python packages spanning LLM APIs, machine learning frameworks, visualization tools, and cloud services, all centralized in [`requirements.txt`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/requirements.txt) for one-command installation.**

SakanaAI/AI-Scientist-v2 is an open-source autonomous research framework that generates scientific papers using large language models. The **dependencies for AI Scientist v2** enable everything from API communication with Anthropic and OpenAI to distributed storage on AWS S3, forming the technical backbone of the automated experimentation pipeline.

## LLM API and Client Libraries

The system interacts with frontier language models through dedicated client libraries wrapped with robust retry mechanisms.

- **`openai`** and **`anthropic`**: Provide native SDK access to GPT-4, Claude, and other frontier models. These are imported and configured in [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) to instantiate chat clients.
- **`backoff`**: Implements exponential backoff logic around API calls to handle rate limits and transient network failures gracefully.
- **`tiktoken`**: Counts tokens in prompts and responses for usage tracking and cost estimation before sending requests to OpenAI endpoints.

## Machine Learning and Data Processing

Core numerical and ML operations rely on industry-standard libraries distributed across the `ai_scientist/treesearch/` modules and experiment runners.

- **`numpy`**: Powers numerical scoring and array operations within tree search algorithms.
- **`transformers`** and **`datasets`**: Enable on-the-fly loading and fine-tuning of pre-trained Hugging Face models during experimentation.
- **`wandb`**: Integrates Weights & Biases for experiment tracking and metric logging.
- **`tqdm`** and **`rich`**: Render live progress bars and styled console output during long-running research loops.
- **`humanize`**: Formats timestamps and byte sizes into human-readable strings for logging.

## Visualization and Document Generation

Publication-ready figures and PDF handling require specialized plotting and document parsing tools utilized in `ai_scientist/perform_plotting` and `ai_scientist/perform_writeup`.

- **`matplotlib`** and **`seaborn`**: Generate line plots, histograms, and statistical visualizations from experiment DataFrames.
- **`pypdf`** and **`pymupdf4llm`**: Parse existing research papers and extract text from PDFs for context injection into LLM prompts.

## Cloud Infrastructure and Storage

Distributed experiments persist artifacts to object storage using AWS SDKs referenced throughout `ai_scientist/utils/`.

- **`boto3`** and **`botocore`**: Interface with AWS S3 to upload model checkpoints, JSON results, and generated LaTeX papers to remote buckets.

## Configuration and Utility Tooling

Supporting infrastructure for code quality, graph processing, and typed configurations appears across the codebase.

- **`omegaconf`**: Manages hierarchical YAML configurations with type safety.
- **`dataclasses-json`**: Serializes experiment state objects to JSON for reproducibility.
- **`python-igraph`**: Implements graph algorithms for the tree-search planning logic in `ai_scientist/treesearch/`.
- **`black`**, **`jsonschema`**, **`genson`**, **`funcy`**, **`coolname`**, and **`shutup`**: Handle code formatting, schema validation, functional programming utilities, random name generation, and suppression of superfluous warnings.

## Architecture Integration

In [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py), the `create_client` function combines `openai`/`anthropic` clients with `backoff` retry wrappers and `tiktoken` accounting. The tree search orchestration in `ai_scientist/treesearch/` leverages `numpy` for numerical operations, `python-igraph` for maintaining search trees, and `rich`/`tqdm` for real-time status displays. Cloud persistence layers in `ai_scientist/utils/` use `boto3` to stream large artifacts to S3 without blocking the main research loop.

## Installation and Setup

Install all dependencies from the repository root using the centralized requirements file:

```bash
pip install -r requirements.txt

```

This command installs the complete dependency graph including transitive requirements for PyTorch, Hugging Face ecosystems, and AWS tooling.

## Practical Implementation Examples

### Initializing an LLM Client with Retry Logic

The `create_client` wrapper in [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) abstracts raw SDK initialization and automatically configures `backoff` retries:

```python
from ai_scientist.llm import create_client

client = create_client(provider="openai", model="gpt-4o-mini")

response = client.chat.complete(
    messages=[{"role": "user", "content": "Explain the significance of the Fourier transform."}]
)
print(response["content"])

```

### Generating Research Visualizations

Experiment metrics convert to publication figures using the visualization stack:

```python
import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

# Assume df contains experiment metrics

sns.lineplot(data=df, x="epoch", y="accuracy", hue="run_id")
plt.title("Training Accuracy Over Epochs")
plt.savefig("results/accuracy_plot.png")

```

### Persisting Artifacts to AWS S3

The cloud storage integration enables durable backup of research outputs:

```python
import boto3
import json

s3 = boto3.client("s3")
bucket = "ai-scientist-artifacts"

payload = json.dumps({"run_id": "abc123", "metrics": {"accuracy": 0.92}})
s3.put_object(Bucket=bucket, Key="runs/abc123/results.json", Body=payload)

```

## Summary

- **AI Scientist v2 dependencies** are declared in [`requirements.txt`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/requirements.txt) and installable via `pip install -r requirements.txt`.
- **LLM interaction** relies on `openai`, `anthropic`, `backoff`, and `tiktoken` for robust API communication with retry logic.
- **Machine learning workflows** use `numpy`, `transformers`, `datasets`, and `wandb` for model handling and experiment tracking.
- **Visualization and document processing** depend on `matplotlib`, `seaborn`, `pypdf`, and `pymupdf4llm` for figure generation and PDF parsing.
- **Cloud persistence** requires `boto3` and `botocore` for S3 storage of research artifacts.
- **Configuration and search algorithms** utilize `omegaconf`, `dataclasses-json`, and `python-igraph` for typed configs and tree-based planning.

## Frequently Asked Questions

### How do I install all dependencies for AI Scientist v2?

Run `pip install -r requirements.txt` from the repository root directory. This installs the complete Python ecosystem including LLM clients, ML frameworks, and AWS SDKs required by [`launch_scientist_bfts.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/launch_scientist_bfts.py) and supporting modules.

### Which packages handle LLM API retries and token counting?

The `backoff` library implements exponential retry logic for flaky API calls, while `tiktoken` performs pre-flight token counting. Both wrap the `openai` and `anthropic` clients inside [`ai_scientist/llm.py`](https://github.com/SakanaAI/AI-Scientist-v2/blob/main/ai_scientist/llm.py) to ensure reliable communication with rate-limited endpoints.

### What visualization libraries does AI Scientist v2 use for research figures?

The system uses `matplotlib` and `seaborn` for statistical plotting and `pypdf`/`pymupdf4llm` for PDF text extraction. These appear in the writeup and plotting modules to generate publication-quality figures and parse existing literature.

### How does AI Scientist v2 manage cloud storage of experiment artifacts?

The `boto3` and `botocore` packages provide AWS S3 integration within `ai_scientist/utils/`, enabling the system to serialize JSON results and upload large model checkpoints to remote storage buckets during distributed experiments.